Skip to content

Add PackageRefSet and SplitRefSet for reflist merging & subtraction - #1617

Draft
refi64 wants to merge 2 commits into
aptly-dev:masterfrom
refi64:wip/refi64/refsets
Draft

Add PackageRefSet and SplitRefSet for reflist merging & subtraction#1617
refi64 wants to merge 2 commits into
aptly-dev:masterfrom
refi64:wip/refi64/refsets

Conversation

@refi64

@refi64 refi64 commented Aug 8, 2026

Copy link
Copy Markdown
Contributor

(draft because it depends on #1606)

Reflist merging uses a simple, linear algorithm, where both reflists are walked in tandem top-to-bottom, placing the lesser reference on the resulting list at each step and removing duplicates. Although this works well for short reflists, it scales poorly when merging many reflists, because the leftmost list is effectively re-scanned every run. This is quadratic time and causes a great deal of memory churn from all the allocations.

A more efficient mechanism appears if we step back and consider the use cases of merging:

  • For package publishing cleanup and database cleanup: we're just trying to merge a large number of lists where refs that are 100% identical (arch, name, version, and hash) are deduplicated. Performance for large numbers of lists here is important, since it will iterate over every published reflist (for publishing cleanup) or every single reflist in the database (for database cleanup).
  • When merging snapshots, and neither -latest nor -no-remove (or API equivalents) are set, overrideMatching=true and ignoreDuplicates=false. In this scenario, in addition to the above deduplication, linear runs of references with a given architecture and name in a later list will override refs in former lists that had the same architecture/name.
  • When creating a new snapshot with multiple sources (API only), or merging snapshots with -no-remove, identical deduplication as in the first case is still performed, but we also want refs that are identical in all but hash (ignoreDuplicates=false) to be overridden by the refs in later lists.
    • This is also used when merging with -latest, since all versions need to be retained in order to identify the latest one via proper version comparisons.

All three of these are actually quite similar; the most significant thing that changes is the "key" used for uniqueness. In the latter two cases, in addition to deduplicating based on the full key, we deduplicate based on "key - {version, hash}" or "key - hash", respectively.

Thus, with these observations, we can create a different approach to merging, the "RefSet", that uses maps. Instead of creating a new reflist for every merge, a single map is created to hold the intermediate state for all future merges. It is set with a key based on the identity we're currently using to deduplicate, pointing to an index into a currently-unsorted slice of refs. As new lists get added, duplicates result in matching runs of refs in the resulting slice to be replaced with the runs in later lists. Then, at the very end, we sort the resulting refs and place them in a reflist.

This brings several advantages:

  • Every reflist added requires iteration relative to the number of refs in only the new reflist, rather than the new one and the old one.
  • Subtracting refsets from reflists is rather cheap, since we just perform a map lookup for each ref in the source list.
  • For split reflists:
    • When using the full ref as the identity, merge order doesn't matter, so we can completely skip buckets whose digest has already been seen.
    • When using a different identity, we can still skip buckets that were seen on the very last AddList call.
    • We can defer bucket digest calculations until the very end, which cuts down on redundant sha256 operations.

The resulting trade-off is that the initial overhead can be higher (around 10x slower, but at these small reflist counts), in particular when you have higher numbers of duplicates and a small number of lists. In the newly-added benchmarks, at the highest duplicate frequency, refsets start out at around 10x slower (though with only 2 lists, the total time is still ~3.7ms), but by 32 lists, they're a bit better at runtime, and nearly 6x better in memory allocated. At 512 lists, refsets are nearly 5x faster (~108ms -> ~22ms) and allocate over 96x(!) less memory.

That is their worst performance in the benchmark. Fewer duplicate counts perform even better, and WithoutHash/WithoutVersion top out at over 51x faster (~1.2s -> ~4.3ms).

In our own database, with ~69k published snapshots, this cuts the publish time from ~33min -> ~3min. I couldn't even wait long enough for a full db cleanup on >200k reflists to finish, but it had made it barely 10% through loading the reflists in over 90 minutes(!). With refsets, the full process finished in under 3 minutes.

In addition to all of this, the new implementation does fix some bugs in the previous one, where certain edge cases when overrideMatching=true would result in refs from an older list being preserved. This can be seen in the first check of the new TestMergingEdgeCases, where the previous implementation would fail because abc 7 would be preserved in the final list.

Checklist

  • unit-test added (if change is algorithm)
  • functional test added/updated (if change is functional)
  • man page updated (if applicable)
  • bash completion updated (if applicable)
  • documentation updated
  • author name in AUTHORS

refi64 added 2 commits July 31, 2026 10:39
In current aptly, each repository and snapshot has its own reflist in
the database. This brings a few problems with it:

- Given a sufficiently large repositories and snapshots, these lists can
  get enormous, reaching >1MB. This is a problem for LevelDB's overall
  performance, as it tends to prefer values around the confiruged block
  size (defaults to just 4KiB).
- When you take these large repositories and snapshot them, you have a
  full, new copy of the reflist, even if only a few packages changed.
  This means that having a lot of snapshots with a few changes causes
  the database to basically be full of largely duplicate reflists.
- All the duplication also means that many of the same refs are being
  loaded repeatedly, which can cause some slowdown but, more notably,
  eats up huge amounts of memory.
- Adding on more and more new repositories and snapshots will cause the
  time and memory spent on things like cleanup and publishing to grow
  roughly linearly.

At the core, there are two problems here:

- Reflists get very big because there are just a lot of packages.
- Different reflists can tend to duplicate much of the same contents.

*Split reflists* aim at solving this by separating reflists into 64
*buckets*. Package refs are sorted into individual buckets according to
the following system:

- Take the first 3 letters of the package name, after dropping a `lib`
  prefix. (Using only the first 3 letters will cause packages with
  similar prefixes to end up in the same bucket, under the assumption
  that packages with similar names tend to be updated together.)
- Take the 64-bit xxhash of these letters. (xxhash was chosen because it
  relatively good distribution across the individual bits, which is
  important for the next step.)
- Use the first 6 bits of the hash (range [0:63]) as an index into the
  buckets.

Once refs are placed in buckets, a sha256 digest of all the refs in the
bucket is taken. These buckets are then stored in the database, split
into roughly block-sized segments, and all the repositories and
snapshots simply store an array of bucket digests.

This approach means that *repositories and snapshots can share their
reflist buckets*. If a snapshot is taken of a repository, it will have
the same contents, so its split reflist will point to the same buckets
as the base repository, and only one copy of each bucket is stored in
the database. When some packages in the repository change, only the
buckets containing those packages will be modified; all the other
buckets will remain unchanged, and thus their contents will still be
shared. Later on, when these reflists are loaded, each bucket is only
loaded once, short-cutting loaded many megabytes of data. In effect,
split reflists are essentially copy-on-write, with only the changed
buckets stored individually.

Changing the disk format means that a migration needs to take place, so
that task is moved into the database cleanup step, which will migrate
reflists over to split reflists, as well as delete any unused reflist
buckets.

All the reflist tests are also changed to additionally test out split
reflists; although the internal logic is all shared (since buckets are,
themselves, just normal reflists), some special additions are needed to
have native versions of the various reflist helper methods.

In our tests, we've observed the following improvements:

- Memory usage during publish and database cleanup, with
  `GOMEMLIMIT=2GiB`, goes down from ~3.2GiB (larger than the memory
  limit!) to ~0.7GiB, a decrease of ~4.5x.
- Database size decreases from 1.3GB to 367MB.

*In my local tests*, publish times had also decreased down to mere
seconds but the same effect wasn't observed on the server, with the
times staying around the same. My suspicions are that this is due to I/O
performance: my local system is an M1 MBP, which almost certainly has
much faster disk speeds than our DigitalOcean block volumes. Split
reflists include a side effect of requiring more random accesses from
reading all the buckets by their keys, so if your random I/O
performance is slower, it might cancel out the benefits. That being
said, even in that case, the memory usage and database size advantages
still persist.

---

Since aptly-dev#1235 and aptly-dev#1282, there were a few improvements:

- I did a fresh rebase on the latest git branch.
- Some strange / problematic edge-case behavior and error handling was
  fixed.
- Tests were improved.

During this time, we've kept split reflists running on our infra, with
no major issues stemming from it. Additionally, the number of snapshots
in the database grew to **over 200k**, with nearly 70k of those actually
published (the reason for this difference is some gaps in our cleanup,
oops). Despite that growth, the total database size is ~1.5GiB, which is
barely larger than our original database was with just a few thousand
snapshots.

Signed-off-by: Ryan Gonzalez <ryan.gonzalez@collabora.com>
Reflist merging uses a simple, linear algorithm, where both reflists are
walked in tandem top-to-bottom, placing the lesser reference on the
resulting list at each step and removing duplicates. Although this works
well for short reflists, it scales poorly when merging *many* reflists,
because the leftmost list is effectively re-scanned every run. This is
quadratic time and causes a great deal of memory churn from all the
allocations.

A more efficient mechanism appears if we step back and consider the use
cases of merging:

- For package publishing cleanup and database cleanup: we're just trying
  to merge a large number of lists where refs that are 100% identical
  (arch, name, version, and hash) are deduplicated. Performance for
  large numbers of lists here is important, since it will iterate over
  *every* published reflist (for publishing cleanup) or *every single
  reflist in the database* (for database cleanup).
- When merging snapshots, and neither `-latest` nor `-no-remove` (or API
  equivalents) are set, `overrideMatching=true` and
  `ignoreDuplicates=false`. In this scenario, in addition to the above
  deduplication, linear runs of references with a given architecture and
  name in a later list will override refs in former lists that had the
  same architecture/name.
- When creating a new snapshot with multiple sources (API only), or
  merging snapshots with `-no-remove`, identical deduplication as in the
  first case is still performed, but we also want refs that are
  identical in all but hash (`ignoreDuplicates=false`) to be overridden
  by the refs in later lists.
  - This is also used when merging with `-latest`, since all versions
    need to be retained in order to identify the latest one via proper
    version comparisons.

All three of these are actually quite similar; the most significant
thing that changes is *the "key" used for uniqueness*. In the latter two
cases, in addition to deduplicating based on the full key, we
deduplicate based on "key - {version, hash}" or "key - hash",
respectively.

Thus, with these observations, we can create a different approach to
merging, the "RefSet", that uses *maps*. Instead of creating a new
reflist for every merge, a single map is created to hold the
intermediate state for all future merges. It is set with a key based on
the identity we're currently using to deduplicate, pointing to an index
into a currently-unsorted slice of refs. As new lists get added,
duplicates result in matching runs of refs in the resulting slice to be
replaced with the runs in later lists. Then, at the very end, we sort
the resulting refs and place them in a reflist.

This brings several advantages:

- Every reflist added requires iteration relative to the number of refs
  in *only* the new reflist, rather than the new one *and* the old one.
- Subtracting refsets from reflists is rather cheap, since we just
  perform a map lookup for each ref in the source list.
- For split reflists:
  - When using the full ref as the identity, merge order doesn't matter,
    so we can completely skip buckets whose digest has already been
    seen.
  - When using a different identity, we can still skip buckets that were
    seen on the very last AddList call.
  - We can defer bucket digest calculations until the very end, which
    cuts down on redundant sha256 operations.

The resulting trade-off is that the initial overhead can be higher
(around 10x slower, but at these small reflist counts), in particular
when you have higher numbers of duplicates and a small number of lists.
In the newly-added benchmarks, at the highest duplicate frequency,
refsets start out at around 10x slower (though with only 2 lists, the
total time is still ~3.7ms), but by 32 lists, they're a bit better at
runtime, and nearly 6x better in memory allocated. At 512 lists, refsets
are nearly 5x faster (~108ms -> ~22ms) and allocate over 96x(!) less
memory.

That is their *worst* performance in the benchmark. Fewer duplicate
counts perform even better, and WithoutHash/WithoutVersion top out at
over 51x faster (~1.2s -> ~4.3ms).

In our own database, with ~69k published snapshots, this cuts the
publish time from ~33min -> ~3min. I couldn't even wait long enough for
a full `db cleanup` on >200k reflists to finish, but it had made it
barely 10% through loading the reflists in over 90 minutes(!). With
refsets, the full process finished in under 3 minutes.

In addition to all of this, the new implementation does *fix* some bugs
in the previous one, where certain edge cases when overrideMatching=true
would result in refs from an older list being preserved. This can be
seen in the first check of the new TestMergingEdgeCases, where the
previous implementation would *fail* because `abc 7` would be preserved
in the final list.
@codecov

codecov Bot commented Aug 8, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 83.98357% with 156 lines in your changes missing coverage. Please review.
✅ Project coverage is 77.26%. Comparing base (f59b0d2) to head (a6fecd2).

Files with missing lines Patch % Lines
deb/reflist.go 86.68% 30 Missing and 21 partials ⚠️
api/db.go 17.02% 35 Missing and 4 partials ⚠️
cmd/db_cleanup.go 67.04% 20 Missing and 9 partials ⚠️
deb/publish.go 79.16% 5 Missing and 5 partials ⚠️
deb/local.go 76.47% 2 Missing and 2 partials ⚠️
deb/remote.go 75.00% 2 Missing and 2 partials ⚠️
deb/snapshot.go 78.94% 2 Missing and 2 partials ⚠️
api/repos.go 88.00% 3 Missing ⚠️
deb/refset.go 97.47% 2 Missing and 1 partial ⚠️
api/snapshot.go 94.87% 1 Missing and 1 partial ⚠️
... and 4 more
Additional details and impacted files
@@            Coverage Diff             @@
##           master    #1617      +/-   ##
==========================================
- Coverage   77.37%   77.26%   -0.12%     
==========================================
  Files         165      166       +1     
  Lines       15747    16348     +601     
==========================================
+ Hits        12185    12631     +446     
- Misses       2356     2479     +123     
- Partials     1206     1238      +32     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant