Skip to content

[WIP] ANNs for clustering at training time - #1938

Open
HowardHuang1 wants to merge 19 commits into
NVIDIA:mainfrom
HowardHuang1:HH-ANNs-for-clustering-at-training
Open

[WIP] ANNs for clustering at training time #1938
HowardHuang1 wants to merge 19 commits into
NVIDIA:mainfrom
HowardHuang1:HH-ANNs-for-clustering-at-training

Conversation

@HowardHuang1

Copy link
Copy Markdown
Contributor

No description provided.

@copy-pr-bot

copy-pr-bot Bot commented Mar 20, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@aamijar aamijar moved this to In Progress in Unstructured Data Processing Mar 20, 2026
@aamijar aamijar added non-breaking Introduces a non-breaking change feature request New feature or request labels Mar 20, 2026
@aamijar

aamijar commented Mar 20, 2026

Copy link
Copy Markdown
Member

/ok to test d26e916

…sed by upstream changes. Add 2 helper functions to recover use_ann_for_extend flag after serialize/deserialize
…rrors Extend_NearestCentroidLookup_Speedup for clarity
…up_Speedup and BM_IVFPQ_Extend_NearestCentroidLookup_Speedup. Also rename use_ann_for_fit to use_ann_for_build_fit to clear up confusion because ANN for nearest cluster assignment can occur in 2 places within build(): (1) in fit and (2) post-fit
…stfit boolean and test cases. Since extend and postfit both use ann on fixed centroids, we factored out common code. Renamed functions in kmeans_balanced.cuh for clarity: assign_nearest_centroid_cagra now more accurately mirrors assign_nearest_centroid_cagra_with_index_reuse.
@HowardHuang1
HowardHuang1 requested review from a team as code owners June 27, 2026 01:53
…weeps

The old Args() list varied N and K independently with no consistent
ratio (2x-312x points-per-cluster across entries), confounding the two
and making the brute-force-vs-CAGRA crossover point ill-defined.

Replace it with two controlled sweeps over the same K range
(1K-1M clusters): RegisterConstantRatioSweep (N = 5 * K, isolates K's
effect at a fixed points-per-cluster ratio) and RegisterFixedNVaryKSweep
(N held at 2,000,000, isolates K's effect at a fixed dataset size).

Add plot_cluster_assignment_bench.py to run the benchmark and plot
brute force vs CAGRA time against K, with automatic crossover detection
and a --side-by-side mode to compare both sweeps at once.
…terval sweeps

Extends the cluster-assignment benchmark with additional controlled
sweeps to fully characterize the brute-force-vs-CAGRA crossover:

- RegisterFixedKVaryNSweep: holds K fixed, varies N, to check whether
  growing N alone (independent of K) can flip the crossover.
- RegisterGridSweep: coarse (N, K) grid for a 2D crossover boundary.
- BM_ClusterAssignment_CAGRA_SearchOnly / RegisterCagraSearchOnlySweep:
  builds the CAGRA index once outside the timed loop, isolating
  steady-state search cost from one-time graph-build cost.
- RegisterDimSensitivitySweep: checks whether the crossover K shifts
  across realistic embedding dimensions (dim was previously untested
  outside 128).
- BM_ClusterAssignment_CAGRA_AmortizedInterval /
  RegisterCagraAmortizedIntervalSweep: empirically measures "build
  once, search `interval` times back-to-back" as a single timed unit,
  instead of projecting it from separately-measured build/search
  numbers, to see how ann_rebuild_interval actually amortizes build
  cost.

Add plot_all_cluster_assignment_benchmarks.py, producing all 8
analysis panels (the two original sweeps, the three new ones, a
speedup-ratio view, a measured build-vs-search breakdown, and
assignment quality vs ann_rebuild_interval from the existing
CUVS_FIT_ANN_REBUILD_TUNING_BENCH) in one figure.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feature request New feature or request non-breaking Introduces a non-breaking change

Projects

Status: In Progress

Development

Successfully merging this pull request may close these issues.

3 participants