Add PackageRefSet and SplitRefSet for reflist merging & subtraction - #1617
Draft
refi64 wants to merge 2 commits into
Draft
Add PackageRefSet and SplitRefSet for reflist merging & subtraction#1617refi64 wants to merge 2 commits into
refi64 wants to merge 2 commits into
Conversation
In current aptly, each repository and snapshot has its own reflist in the database. This brings a few problems with it: - Given a sufficiently large repositories and snapshots, these lists can get enormous, reaching >1MB. This is a problem for LevelDB's overall performance, as it tends to prefer values around the confiruged block size (defaults to just 4KiB). - When you take these large repositories and snapshot them, you have a full, new copy of the reflist, even if only a few packages changed. This means that having a lot of snapshots with a few changes causes the database to basically be full of largely duplicate reflists. - All the duplication also means that many of the same refs are being loaded repeatedly, which can cause some slowdown but, more notably, eats up huge amounts of memory. - Adding on more and more new repositories and snapshots will cause the time and memory spent on things like cleanup and publishing to grow roughly linearly. At the core, there are two problems here: - Reflists get very big because there are just a lot of packages. - Different reflists can tend to duplicate much of the same contents. *Split reflists* aim at solving this by separating reflists into 64 *buckets*. Package refs are sorted into individual buckets according to the following system: - Take the first 3 letters of the package name, after dropping a `lib` prefix. (Using only the first 3 letters will cause packages with similar prefixes to end up in the same bucket, under the assumption that packages with similar names tend to be updated together.) - Take the 64-bit xxhash of these letters. (xxhash was chosen because it relatively good distribution across the individual bits, which is important for the next step.) - Use the first 6 bits of the hash (range [0:63]) as an index into the buckets. Once refs are placed in buckets, a sha256 digest of all the refs in the bucket is taken. These buckets are then stored in the database, split into roughly block-sized segments, and all the repositories and snapshots simply store an array of bucket digests. This approach means that *repositories and snapshots can share their reflist buckets*. If a snapshot is taken of a repository, it will have the same contents, so its split reflist will point to the same buckets as the base repository, and only one copy of each bucket is stored in the database. When some packages in the repository change, only the buckets containing those packages will be modified; all the other buckets will remain unchanged, and thus their contents will still be shared. Later on, when these reflists are loaded, each bucket is only loaded once, short-cutting loaded many megabytes of data. In effect, split reflists are essentially copy-on-write, with only the changed buckets stored individually. Changing the disk format means that a migration needs to take place, so that task is moved into the database cleanup step, which will migrate reflists over to split reflists, as well as delete any unused reflist buckets. All the reflist tests are also changed to additionally test out split reflists; although the internal logic is all shared (since buckets are, themselves, just normal reflists), some special additions are needed to have native versions of the various reflist helper methods. In our tests, we've observed the following improvements: - Memory usage during publish and database cleanup, with `GOMEMLIMIT=2GiB`, goes down from ~3.2GiB (larger than the memory limit!) to ~0.7GiB, a decrease of ~4.5x. - Database size decreases from 1.3GB to 367MB. *In my local tests*, publish times had also decreased down to mere seconds but the same effect wasn't observed on the server, with the times staying around the same. My suspicions are that this is due to I/O performance: my local system is an M1 MBP, which almost certainly has much faster disk speeds than our DigitalOcean block volumes. Split reflists include a side effect of requiring more random accesses from reading all the buckets by their keys, so if your random I/O performance is slower, it might cancel out the benefits. That being said, even in that case, the memory usage and database size advantages still persist. --- Since aptly-dev#1235 and aptly-dev#1282, there were a few improvements: - I did a fresh rebase on the latest git branch. - Some strange / problematic edge-case behavior and error handling was fixed. - Tests were improved. During this time, we've kept split reflists running on our infra, with no major issues stemming from it. Additionally, the number of snapshots in the database grew to **over 200k**, with nearly 70k of those actually published (the reason for this difference is some gaps in our cleanup, oops). Despite that growth, the total database size is ~1.5GiB, which is barely larger than our original database was with just a few thousand snapshots. Signed-off-by: Ryan Gonzalez <ryan.gonzalez@collabora.com>
Reflist merging uses a simple, linear algorithm, where both reflists are
walked in tandem top-to-bottom, placing the lesser reference on the
resulting list at each step and removing duplicates. Although this works
well for short reflists, it scales poorly when merging *many* reflists,
because the leftmost list is effectively re-scanned every run. This is
quadratic time and causes a great deal of memory churn from all the
allocations.
A more efficient mechanism appears if we step back and consider the use
cases of merging:
- For package publishing cleanup and database cleanup: we're just trying
to merge a large number of lists where refs that are 100% identical
(arch, name, version, and hash) are deduplicated. Performance for
large numbers of lists here is important, since it will iterate over
*every* published reflist (for publishing cleanup) or *every single
reflist in the database* (for database cleanup).
- When merging snapshots, and neither `-latest` nor `-no-remove` (or API
equivalents) are set, `overrideMatching=true` and
`ignoreDuplicates=false`. In this scenario, in addition to the above
deduplication, linear runs of references with a given architecture and
name in a later list will override refs in former lists that had the
same architecture/name.
- When creating a new snapshot with multiple sources (API only), or
merging snapshots with `-no-remove`, identical deduplication as in the
first case is still performed, but we also want refs that are
identical in all but hash (`ignoreDuplicates=false`) to be overridden
by the refs in later lists.
- This is also used when merging with `-latest`, since all versions
need to be retained in order to identify the latest one via proper
version comparisons.
All three of these are actually quite similar; the most significant
thing that changes is *the "key" used for uniqueness*. In the latter two
cases, in addition to deduplicating based on the full key, we
deduplicate based on "key - {version, hash}" or "key - hash",
respectively.
Thus, with these observations, we can create a different approach to
merging, the "RefSet", that uses *maps*. Instead of creating a new
reflist for every merge, a single map is created to hold the
intermediate state for all future merges. It is set with a key based on
the identity we're currently using to deduplicate, pointing to an index
into a currently-unsorted slice of refs. As new lists get added,
duplicates result in matching runs of refs in the resulting slice to be
replaced with the runs in later lists. Then, at the very end, we sort
the resulting refs and place them in a reflist.
This brings several advantages:
- Every reflist added requires iteration relative to the number of refs
in *only* the new reflist, rather than the new one *and* the old one.
- Subtracting refsets from reflists is rather cheap, since we just
perform a map lookup for each ref in the source list.
- For split reflists:
- When using the full ref as the identity, merge order doesn't matter,
so we can completely skip buckets whose digest has already been
seen.
- When using a different identity, we can still skip buckets that were
seen on the very last AddList call.
- We can defer bucket digest calculations until the very end, which
cuts down on redundant sha256 operations.
The resulting trade-off is that the initial overhead can be higher
(around 10x slower, but at these small reflist counts), in particular
when you have higher numbers of duplicates and a small number of lists.
In the newly-added benchmarks, at the highest duplicate frequency,
refsets start out at around 10x slower (though with only 2 lists, the
total time is still ~3.7ms), but by 32 lists, they're a bit better at
runtime, and nearly 6x better in memory allocated. At 512 lists, refsets
are nearly 5x faster (~108ms -> ~22ms) and allocate over 96x(!) less
memory.
That is their *worst* performance in the benchmark. Fewer duplicate
counts perform even better, and WithoutHash/WithoutVersion top out at
over 51x faster (~1.2s -> ~4.3ms).
In our own database, with ~69k published snapshots, this cuts the
publish time from ~33min -> ~3min. I couldn't even wait long enough for
a full `db cleanup` on >200k reflists to finish, but it had made it
barely 10% through loading the reflists in over 90 minutes(!). With
refsets, the full process finished in under 3 minutes.
In addition to all of this, the new implementation does *fix* some bugs
in the previous one, where certain edge cases when overrideMatching=true
would result in refs from an older list being preserved. This can be
seen in the first check of the new TestMergingEdgeCases, where the
previous implementation would *fail* because `abc 7` would be preserved
in the final list.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #1617 +/- ##
==========================================
- Coverage 77.37% 77.26% -0.12%
==========================================
Files 165 166 +1
Lines 15747 16348 +601
==========================================
+ Hits 12185 12631 +446
- Misses 2356 2479 +123
- Partials 1206 1238 +32 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
(draft because it depends on #1606)
Reflist merging uses a simple, linear algorithm, where both reflists are walked in tandem top-to-bottom, placing the lesser reference on the resulting list at each step and removing duplicates. Although this works well for short reflists, it scales poorly when merging many reflists, because the leftmost list is effectively re-scanned every run. This is quadratic time and causes a great deal of memory churn from all the allocations.
A more efficient mechanism appears if we step back and consider the use cases of merging:
-latestnor-no-remove(or API equivalents) are set,overrideMatching=trueandignoreDuplicates=false. In this scenario, in addition to the above deduplication, linear runs of references with a given architecture and name in a later list will override refs in former lists that had the same architecture/name.-no-remove, identical deduplication as in the first case is still performed, but we also want refs that are identical in all but hash (ignoreDuplicates=false) to be overridden by the refs in later lists.-latest, since all versions need to be retained in order to identify the latest one via proper version comparisons.All three of these are actually quite similar; the most significant thing that changes is the "key" used for uniqueness. In the latter two cases, in addition to deduplicating based on the full key, we deduplicate based on "key - {version, hash}" or "key - hash", respectively.
Thus, with these observations, we can create a different approach to merging, the "RefSet", that uses maps. Instead of creating a new reflist for every merge, a single map is created to hold the intermediate state for all future merges. It is set with a key based on the identity we're currently using to deduplicate, pointing to an index into a currently-unsorted slice of refs. As new lists get added, duplicates result in matching runs of refs in the resulting slice to be replaced with the runs in later lists. Then, at the very end, we sort the resulting refs and place them in a reflist.
This brings several advantages:
The resulting trade-off is that the initial overhead can be higher (around 10x slower, but at these small reflist counts), in particular when you have higher numbers of duplicates and a small number of lists. In the newly-added benchmarks, at the highest duplicate frequency, refsets start out at around 10x slower (though with only 2 lists, the total time is still ~3.7ms), but by 32 lists, they're a bit better at runtime, and nearly 6x better in memory allocated. At 512 lists, refsets are nearly 5x faster (~108ms -> ~22ms) and allocate over 96x(!) less memory.
That is their worst performance in the benchmark. Fewer duplicate counts perform even better, and WithoutHash/WithoutVersion top out at over 51x faster (~1.2s -> ~4.3ms).
In our own database, with ~69k published snapshots, this cuts the publish time from ~33min -> ~3min. I couldn't even wait long enough for a full
db cleanupon >200k reflists to finish, but it had made it barely 10% through loading the reflists in over 90 minutes(!). With refsets, the full process finished in under 3 minutes.In addition to all of this, the new implementation does fix some bugs in the previous one, where certain edge cases when overrideMatching=true would result in refs from an older list being preserved. This can be seen in the first check of the new TestMergingEdgeCases, where the previous implementation would fail because
abc 7would be preserved in the final list.Checklist
AUTHORS