On this page· 8

Merge packs

  • file format 6
  • Rust SDK 1.0 preview

Goal: make one pack from several, and know which record is kept when two packs disagree.

Merge, sometimes called stacking, takes one or more packs and a strategy, and returns one pack and a report.

Two packs, one conflict

  1. 1. Collect the packs as byte slices.
  2. 2. Call merge(&[&a, &b], &MergeOptions::default()).
  3. 3. Use the returned bytes as the new pack, and read the report.

Needs the Rust SDK

merge.rsrust
// Merge two packs and choose how a conflict between two versions of
// one entity is settled.
use plxi_sdk::{
    merge, write, MergeOptions, MergeStrategy, OpenOptions, Pack, Record,
};

type Result<T> = std::result::Result<T, plxi_sdk::Error>;

fn pack(lines: &[&str]) -> Result<Vec<u8>> {
    let records: Vec<Record> = lines
        .iter()
        .map(|l| Record::from_json_line(l))
        .collect::<Result<_>>()?;
    write(&records, None)
}

fn name_of_e1(bytes: Vec<u8>) -> Result<Option<String>> {
    let pack = Pack::from_bytes(bytes, OpenOptions::default())?;
    for record in pack.records() {
        let record = record?;
        if let Some(e) = record.as_entity() {
            if e.id == "e1" {
                return Ok(e.name.map(str::to_string));
            }
        }
    }
    Ok(None)
}

fn main() -> Result<()> {
    let a = pack(&[
        r#"{"k":"ent","id":"e1","t":"Person","salience":0.5}"#,
        r#"{"k":"ent","id":"e2","t":"Person"}"#,
        r#"{"k":"rel","src":"e1","tgt":"e2","kind":"knows"}"#,
    ])?;
    let b = pack(&[
        r#"{"k":"ent","id":"e1","t":"Person","salience":0.75,"name":"B"}"#,
        r#"{"k":"ent","id":"e3","t":"Person"}"#,
        r#"{"k":"rel","src":"e1","tgt":"e3","kind":"knows"}"#,
    ])?;

    // Default strategy: the incoming entity wins only with a strictly
    // greater salience. 0.75 > 0.5, so pack b's e1 is taken.
    let default = MergeOptions::default();
    assert_eq!(default.strategy, MergeStrategy::HigherSalience);
    let (merged, report) = merge(&[&a, &b], &default)?;
    assert_eq!((report.records_inserted, report.conflicts), (5, 1));
    assert_eq!(name_of_e1(merged)?, Some("B".to_string()));
    println!("{}", report.to_json());

    // `Union` keeps the record that was there first.
    let union = MergeOptions::with_strategy(MergeStrategy::Union);
    let (merged, _) = merge(&[&a, &b], &union)?;
    assert_eq!(name_of_e1(merged)?, None);

    // Strategies have stable names, for configuration files.
    assert_eq!(MergeStrategy::Latest.as_str(), "latest");
    let manual = MergeStrategy::from_name("manual");
    assert_eq!(manual, Some(MergeStrategy::Manual));
    assert_eq!(MergeStrategy::from_name("newest"), None);
    Ok(())
}
Output
{"records_inserted":5,"records_updated":1,"records_skipped":0,"stubs_resolved":0,"conflicts":1,"sections_merged":0,"sections_deduped":0}

Ran with cargo run --example 05_merge · exit status 0

Both packs have an entity e1 with different members. That is a conflict. The default strategy keeps pack b's e1, because its salience, 0.75, is greater than 0.5.

The record rules

Inputs are taken in argument order, one record at a time. Each record is looked up by its dedup key, (kind, primary id).

Table 1. What merge does with an incoming record: 6 cases
CaseResultReport counters
the key is newinsertedrecords_inserted
an equal record is already thereskippedrecords_skipped
both are ent; the one there is a stub and the incoming one is notthe incoming record replaces the stub, under every strategystubs_resolved records_updated
both are ent; the incoming one is a stub and the one there is notthe one there is kept, under every strategyrecords_skipped
both are ent, neither is a stub, and they differa conflict, settled by the strategyconflicts records_updated records_skipped
any other kind, and the records differlatest takes the incoming record; every other strategy keeps the one thererecords_updated or records_skipped; not conflicts

Merge drops the payload records of its inputs and writes one new one, so its output always holds exactly one payload record.

The four strategies

Table 2. Merge strategies: 4
NameRustAn entity conflict is settled by
higher_salienceMergeStrategy::HigherSaliencethe incoming entity wins only if its salience is strictly greater; a missing salience counts as 0.0
latestMergeStrategy::Latestthe incoming entity wins
unionMergeStrategy::Unionthe entity already there is kept
manualMergeStrategy::Manualthe entity already there is kept
Table 3. The same two packs in both orders: salience of e1 in the output
Strategymerge(a, b)merge(b, a)
higher_salience0.750.75
latest0.750.5
union0.50.75
manual0.50.75
  • Argument order matters under a conflict. Only higher_salience, with two different salience values, gives the same entity in both orders.
  • latest compares no timestamps. "Latest" means later in the argument list.
  • union keeps one record per key. It does not keep both versions.
  • Through this SDK, manual and union give the same bytes and the same report. The report holds counters, not a list of conflicts.
  • Strategy names are fixed strings. MergeStrategy::from_name accepts higher_salience, latest, union and manual, and returns None for any other name.

The four strategies are conformance cases on plxi.org. Merge cases

Three packs, one stub

Pack a mentions water only as a stub. Pack b has the full water record. Pack c repeats pack a's tea. Merged with union, the stub gives way to the full record and the repeated tea is skipped.

Table 4. Input packs: 3 packs, 9 records
PackRecord
a{"k":"ent","id":"tea","t":"Topic","name":"Tea"}
a{"k":"ent","id":"water","t":"Topic","stub":true}
a{"k":"rel","src":"tea","tgt":"water","kind":"needs"}
b{"k":"ent","id":"water","t":"Topic","name":"Water"}
b{"k":"ent","id":"kettle","t":"Tool","name":"Kettle"}
b{"k":"rel","src":"kettle","tgt":"water","kind":"heats"}
c{"k":"ent","id":"tea","t":"Topic","name":"Tea"}
c{"k":"ent","id":"cup","t":"Tool","name":"Cup"}
c{"k":"rel","src":"cup","tgt":"tea","kind":"holds"}
Table 5. Output: 8 records
Record
{"k":"payload","v":1,"embedded":false,"csdt_file_checksum":0,"sections":[],"refs":[]}
{"k":"ent","v":1,"id":"cup","t":"Tool","name":"Cup","stub":false}
{"k":"ent","v":1,"id":"kettle","t":"Tool","name":"Kettle","stub":false}
{"k":"ent","v":1,"id":"tea","t":"Topic","name":"Tea","stub":false}
{"k":"ent","v":1,"id":"water","t":"Topic","name":"Water","stub":false}
{"k":"rel","v":1,"src":"cup","tgt":"tea","kind":"holds"}
{"k":"rel","v":1,"src":"kettle","tgt":"water","kind":"heats"}
{"k":"rel","v":1,"src":"tea","tgt":"water","kind":"needs"}

Merged as a, b, c and as c, b, a, the output is the same file, with digest 86aff647…1557b. The two reports differ, because a report counts what happened along the way. In one order a stub was replaced. In the other the full record came first, and the stub was skipped.

Packs with binary data

When the inputs have appendices, their sections are placed side by side in the output and exact duplicates are removed. Two Compact sections stay two sections. Each local binding has its section index rewritten to the section's new position; its row index does not change.

Table 6. Bindings before and after: 3
EntityBefore: section, rowAfter: section, row
p10, 00, 0
q10, 01, 0
q20, 11, 1

Verify on the output counts 3 bindings, 3 resolved, 0 dangling.

What merge refuses

Table 7. Merge refusals: 4
InputError kind
no packs at allmerge_conflict
a pack that fails verify steps 1 to 8that step's kind, for example checksum_mismatch
a pack whose appendix holds a tensormerge_conflict
a pack whose appendix uses an older container versionappendix_legacy_version

Merging a single good pack is allowed and is not a no-op: the output has one more record, the payload record. Merging that output again changes nothing.

Change logs are not merged

A pack of mut records is an ordered log, and the records of two logs share sequence numbers. Merge keeps one record per key and raises no error, so one of two changes numbered 1 is lost. Change records

Reference

Sections