Home / Notes / Sixty days, eight H100s: building Version 1
Notes
What NVIDIA's Innovation Lab made possible.
Before Curator Research was incorporated, we completed Version 1 during a 60-day NVIDIA DGX Cloud Innovation Lab allocation, accessed through Inception membership and Brev. The work brought document parsing, retrieval, reranking, model fine-tuning and evaluation together on one compute node. The methods and trained artefacts then moved into our own infrastructure.
22 September 2026 · Curator Research · Content reviewed 29 September 2026
01
The hard part was understanding the source
An answer to a regulatory or legal question needs to lead back to the exact paragraph or chapter, and to the dated circular, direction or judgment that applies. The source material rarely arrives ready for that. We had to collect changing documents, preserve their lineage, read scans and bilingual pages, reconstruct tables that run across pages, and distinguish definitions and exceptions from hard rules, limits, permissions and obligations. Those distinctions have to survive extraction so retrieval and answer generation can cite the right text. Our smaller machines could test individual stages, but could not sustain parsing, enrichment, retrieval and evaluation together. Some difficult PDFs stalled whole batches; faster extraction sometimes lost reading order, tables or page provenance. We began profiling documents by type and measuring both throughput and structural fidelity, then tuned the parsing workers, OCR and table routes around the failure patterns. This gave the later models a more reliable representation of the source to reason over.
02
Turning one machine into an orchestrated pipeline
The grant node supplied eight 80 GB H100s, 208 CPU cores, 1.8 TiB of RAM and about 18.6 TB of local storage. At first we spread document-parsing workers across the GPUs and used NVIDIA Multi-Process Service to share them efficiently. As workloads changed, we tested a split with six GPUs serving three two-GPU Llama-3.3-70B NIM replicas, one GPU continuing parsing and one serving embeddings. Training, reranking and evaluation needed the same machine at other times. Static assignments soon gave way to a pipeline orchestrator with distinct parsing, enrichment and embedding lanes and resource admission based on available CPU, memory, GPU memory and database capacity. We adjusted batch sizes, worker concurrency and GPU placements as bottlenecks moved, and introduced recovery paths for slow documents and interrupted runs. When a late training job slowed, the limiting resource was CPU time feeding the GPUs, so we reduced ingestion pressure instead of assigning more cards. The grant application began with 261,967 documents in scope; the source collection ultimately grew into an approximately 15-million-document catalogue spanning regulatory material, statutes, and Supreme Court and High Court cases. Parsing, enrichment and indexing progressed through separate, traceable stages.
03
Models and retrieval trained for different jobs
We fine-tuned Qwen3-8B and Qwen3.6-35B-A3B variants for enrichment and orchestration, including reasoning-on and direct-response modes. Enrichment extracts structured rules, entities and citations; orchestration plans work across tools. We fine-tuned Nemotron-H 30B with NVIDIA NeMo AutoModel. Llama-3.3-70B served through NVIDIA NIM powered enrichment and supplied teacher targets and offline judgments for training and evaluation. We also trained 8B synthesis adapters for grounded answers. Run manifests, dataset hashes, resumable checkpoints and evaluations of schema fidelity and source grounding kept the training decisions traceable. Retrieval needed its own training. We developed a six-layer MiniLM multisource reranker to order passages after the first search, then compared reranking approaches on an internal set of 439 questions across seven source families. The work led us to strengthen hard-negative construction and evaluation splits as well as the model itself.
04
The node was used, not just reserved
Across measured multi-day runs, mean utilisation was 77.0%; peak memory use reached 627 of 633 GiB, and median power draw above 90% utilisation was 610 W per GPU. The node let us run document processing, model serving and training experiments at a scale our smaller machines could not sustain together.
05
What kept working after the grant
The allocation had an end date, so portability mattered from the start. Quantisation work produced an approximately 6.1 GB Qwen3-8B package and a 21.8 GB W4A16 weight artefact for the 35B mixture-of-experts model. More important than the files' sizes was what we could actually run: a fine-tuned 8B model served on AWS L40S with four adapter identities recorded in the serving configuration, and the 8B multi-adapter stack also passed an end-to-end smoke test on an 8 GB GPU using a public AWQ base. Fine-tuned 35B enrichment and orchestration models ran on AWS GPU and spot machines for defined regulatory tasks. On those tasks, their quality came close to Llama while using resources more efficiently. That work gave us a practical division of labour: 8B adapters for lighter specialist jobs and the 35B path for demanding enrichment. The grant supplied the training and evaluation window; our own machines showed how those results could continue as operating components rather than remain a one-off demonstration.
06
Why Version 2 is about reconciliation
Version 1 demonstrated that parsing, enrichment, retrieval and specialised models could work together. It also exposed a deeper systems problem: a source document and its derived sections, chunks, citations, embeddings and graph relationships live across relational, vector and graph databases. When a source changes or a stage is rerun, those representations can drift apart even while each store appears healthy. Curator Research Version 2 is being built around ongoing, native reconciliation across those stores. The system compares the expected source lineage and derived records with what is actually present, then routes missing or stale work for verification and repair. That is how a researched answer can retain its exact source, effective date and citation path as the knowledge base changes.
07
Beyond the allocation
The work continued after the grant node closed. We moved pipeline and model serving to AWS, used fine-tuned 35B adapters for regulatory enrichment, and kept refining the balance between source quality, retrieval, model choice and operating cost. Version 1 gave us the technical foundation and the confidence to establish Curator Research as a company pursuing document intelligence, connected knowledge, research and governed action.
08
Version 2: useful at different scales
For small professional teams, that means pursuing efficient models and local paths within modest budgets. For larger institutions, it means controlled deployment within tenant boundaries, with sources, dates and execution records that remain reviewable. NVIDIA's Innovation Lab and Brev enabled the concentrated Version 1 work; Version 2 is our work to make those foundations dependable at different scales. Current product stages and available engagements are on our Capabilities page.
Disagree with us? Tell us why.
We publish the measurement, its date and its scope so you can check our reasoning. If you think we got something wrong, write to us.