
Local AI 101: what your own machine can run
A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read MoreBuilding upon the success of Genesis I, we introduce QVAC Genesis II, a major expansion that adds new domains and a total of 148 billion tokens.

Building upon the success of QVAC Genesis I [1] , the largest publicly available synthetic dataset for educational content (41 billion tokens), Tether Data, S.A. de C.V. (Tether Data, we, us, our) introduces QVAC Genesis II, a major expansion that adds 10 new educational domains, 107 billion new tokens, and introduces a new Option-Level Reasoning data generation method. Combined with Genesis I, the dataset now totals 148 billion tokens.
Genesis I focused on core STEM disciplines (Mathematics, Physics, Biology, and Medicine) and demonstrated superior performance compared to existing synthetic datasets like Cosmopedia [2]. Genesis II extends this foundation by:
QVAC Genesis II expands upon Genesis I with the following contributions:
Note: For more detailed information about the base methodology used in Genesis II, including seed data acquisition and quality filtering processes, original prompt templates for Scaling QA, MCQ Answer, LLM-as-a-Judge extraction, and Failure Analysis, pipeline orchestration details (distilabel, vLLM), and model architectures and configurations, please refer . to the comprehensive Genesis I Appendix.
Genesis II builds upon the proven "Learning from Failures" method from Genesis I while introducing a complementary data generation method. This dual-approach methodology maximizes the value extracted from every generated question, whether the model answers correctly or incorrectly.
For complete details on “Learning from Failures” please refer to QVAC Genesis I.

Figure 1. The enhanced Genesis II pipeline: Seed Data → Quality Filter → Scaling QA (generate 4 MCQs per seed) → Model Answering → Compare to Gold Label → Two methods:
Genesis II introduces the Option-Level Reasoning Analysis method, applied to questions that the model answers correctly during the Model Answering phase. While the original Failure Analysis method focused on extracting educational value from model errors, Option-Level Reasoning Analysis ensures that correctly answered questions also contribute high-quality educational content.
Rationale: A model answering a question correctly demonstrates understanding, but the reasoning behind that understanding, along with the explicit explanation of why other options are incorrect, provides valuable educational content. This approach:
Four Output Styles: Similar to Failure Analysis, Option-Level Reasoning Analysis generates educational content in four distinct styles:
Each style analyzes the correct answer option first with detailed reasoning, then systematically examines each incorrect option. The complete prompt templates are documented in the Appendix.
For Genesis II, we expanded the 9 domains from Genesis I to include 10 new educational domains:
New Domains:
The rigorous seed data acquisition (using FineFineWeb [4]), quality filtering (using Ultra-FineWeb-classifier [5]), and prompt engineering methodology from Genesis I is applied to our new domains. For full methodology, refer to Genesis I Blog.
Training a 1.7B parameter model from scratch on 64 GPUs sounds straightforward on paper, until you confront the fragmented landscape of distributed training frameworks.
On one hand, we have HuggingFace Transformers: the standard for model definitions, with thousands of architectures, clean APIs, and a large community. Qwen3-1.7B exists here, complete with its attention patterns, RoPE embeddings, and SwiGLU activations.
On the other hand, we have Megatron-Core: NVIDIA's framework for large-scale training, with optimized CUDA kernels, mature tensor parallelism, and communication patterns refined over years of training at scale. This is where training needs to happen if you want reasonable throughput on 64 GPUs.
The problem is these two worlds don't speak the same language.
The traditional path to using Megatron-Core required rewriting your model from scratch in Megatron's internal format: manually implementing each layer type with the correct parallelism layouts, debugging distributed deadlocks, writing checkpoint conversion scripts, and implementing data loading in Megatron's binary format. This could easily be a multi-month project.
Megatron-Bridge solves this by automatically converting HuggingFace model definitions into Megatron-compatible formats. This lets us use the Qwen3-1.7B architecture without rewriting it:
All three models were trained separately on a 64-GPU cluster (8 nodes with 8 NVIDIA H100 GPUs each), connected via InfiniBand.
Component | Specification |
|---|---|
GPUs | 64 × NVIDIA H100 (80GB) |
Nodes | 8 nodes × 8 GPUs each |
Interconnect | InfinityBand with GPU Direct RDMA |
Container | NVIDIA NeMo 25.09 |
Distributing the training across 64 GPUs requires deciding how to split the work. We use a combination of tensor parallelism and data parallelism:
Parallelism Type | Size | What It Does |
|---|---|---|
Tensor (TP) | 2 | Splits attention and feed-forward layers across 2 GPUs |
Pipeline (PP) | 1 | No pipeline splitting (model fits in memory) |
Data (DP) | 32 | 32 parallel workers process different batches |
Why TP=2? Tensor parallelism requires frequent communication between GPUs. At TP=2, this communication stays within a single node using fast NVLink. Higher TP would require cross-node communication on every layer, reducing throughput.
Why PP=1? Pipeline parallelism is useful for very large models that don't fit in memory. At 1.7B parameters with TP=2, the model fits comfortably, so pipeline splitting would only add overhead.
Why DP=32? After allocating GPUs for tensor parallelism (64 ÷ 2 = 32), the remaining capacity goes to data parallelism. Each of the 32 workers processes different batches in parallel, then synchronizes gradients.
The batch size configuration balances memory constraints with training efficiency:
Parameter | Value | Rationale |
|---|---|---|
Micro Batch Size | 4 per GPU | Limited by GPU memory with 4,096 token sequences |
Gradient Accumulation | 16 steps | Accumulate gradients before synchronizing |
Global Batch Size | 2,048 sequences | 4 × 32 workers × 16 accumulation steps |
Tokens per Step | ~8.4M | 2,048 sequences × 4,096 tokens |
The micro batch size of 4 might seem small for 80GB GPUs, but at 4,096 tokens per sequence with tensor parallelism, this is near the memory limit. We compensate by accumulating gradients over 16 forward passes before updating weights, reaching our target global batch size of 2,048 sequences (~8.4 million tokens per training step).
Rather than training a single model, we designed an experiment to evaluate how synthetic data, created using multiple prompts, generalizes. We trained three distinct models from scratch to create a rigorous comparison:
To ensure a fair comparison, all three models utilized the same hyperparameters and compute budget scaled to token count, differing only in the data composition.
Hyperparameters
We standardized the training duration for all Genesis II runs to a single epoch, ensuring the model encountered each unique synthetic example only once. This strategy mitigates the memorization of repeated tokens and encourages the learning of underlying logic.
Learning rate schedule: We start with a warmup period (10% of the epoch) where the learning rate gradually increases to 2×10⁻⁴. This helps stabilize early training when the randomly initialized model produces noisy gradients. After warmup, the learning rate follows a cosine decay down to 2×10⁻⁵, allowing the model to settle into better solutions as it converges toward the end of the epoch.
BF16 precision: We use bfloat16 mixed precision with Flash Attention 2. This configuration significantly reduces memory usage and speeds up training throughput on the H100s without any meaningful loss in convergence accuracy compared to FP32.
Transforming raw text into a format that Megatron-Core can efficiently ingest requires a multi-stage pipeline. This preprocessing happens once before training begins - meaning any mistakes here propagate through the entire run.
Stage 1: Concatenation and Filtering
Our Genesis II data comprises thousands of individual JSONL files produced by the data generation workers. The first step consolidates these into a single file while applying quality filters:
This filtering step is important because even small amounts of low-quality data (empty documents, truncated text, failed generations) can degrade training.
Stage 2: Tokenization and Binary Conversion
The filtered JSONL is then processed by Megatron's preprocessing tool, which:
The result is two files: a .bin file containing all token IDs packed end-to-end, and an .idx file containing the byte offsets for each document. This format allows the data loader to seek directly to any document without reading the entire file.
Stage 3: Sequence Packing During Training
During training, Megatron's data loader constructs training sequences from this binary format:
This "packing" approach is more efficient than padding each document to a fixed length so that short documents don't waste compute on padding tokens, and the model naturally learns to handle document transitions. A single training sequence might contain 2-3 documents if they're short enough.
Genesis I Domains:
Domain | Number of Samples | No of Tokens (in B) |
|---|---|---|
High school biology | 3,818,070 | 4.511 |
College biology | 3,286,648 | 3.927 |
Professional medicine | 1,552,474 | 1.884 |
College medicine | 5,164,247 | 6.218 |
High school mathematics | 3,244,240 | 4.277 |
College mathematics | 5,895,052 | 8.243 |
High school physics | 2,277,880 | 3.061 |
College physics | 4,281,062 | 5.814 |
Conceptual physics | 2,354,184 | 2.973 |
Genesis I Total | 31,873,857 | 40.906 |
Genesis II New Domains — Failure Analysis Data:
Domain | Number of Samples | No of Tokens (in B) |
|---|---|---|
College physics | 4,144,798 | 6.24 |
Astronomy | 4,716,117 | 6.21 |
Econometrics | 3,486,501 | 5.24 |
College chemistry | 3,964,112 | 5.07 |
Electrical Engineering | 3,901,901 | 4.96 |
College computer science | 3,889,696 | 4.77 |
Geography | 3,992,646 | 4.60 |
High school statistics | 3,354,353 | 4.47 |
High school chemistry | 3,327,350 | 4.15 |
High school computer science | 3,365,258 | 4.06 |
Machine learning | 3,133,569 | 3.87 |
Failure Analysis Total | 41,276,301 | 53.64 |
Genesis II New Domains — Option-Level Reasoning Analysis Data:
Domain | Number of Samples | No of Tokens (in B) |
|---|---|---|
Machine learning | 4,636,066 | 5.51 |
High school statistics | 4,424,565 | 5.41 |
High school chemistry | 4,464,847 | 5.22 |
Econometrics | 3,871,249 | 5.06 |
College chemistry | 4,182,669 | 5.01 |
College physics | 3,672,394 | 4.81 |
Geography | 4,301,699 | 4.77 |
Astronomy | 3,970,849 | 4.71 |
College computer science | 3,851,555 | 4.47 |
Electrical Engineering | 3,758,536 | 4.44 |
High school computer science | 3,885,236 | 4.41 |
Option-Level Reasoning Analysis Total | 45,019,665 | 53.82 |
Genesis II Combined Total (Failure + Option-Level Reasoning): 86,295,966 samples | 107.46B tokens
Combined Genesis I + Genesis II Total: 118,169,823 samples | 148.37B tokens
Genesis II introduces an evaluation framework that goes beyond simple accuracy. We evaluate models using LLM-as-a-Judge via the OpenCompass framework. The judge analyzes each model response and classifies it into one of the following categories:
When the LLM judge evaluates a model's response, it determines whether the response contains a clear, extractable answer:
✓ Valid Answers — The judge successfully identifies a single, clear answer in the response:
✗ Invalid Answers — The judge cannot extract a valid answer from the response. This includes two types:
Based on the judge's classification, we compute the following metrics:
Metric | Definition |
|---|---|
Valid Answer Rate | Percentage of responses where the judge identified a clear, single answer |
No Answer Rate | Percentage of responses where the judge found no clear answer |
Multiple Answers Rate | Percentage of responses with multiple conflicting answers |
Accuracy | Percentage of valid answers that are correct |
Formula: Valid Answer Rate = 100% - No Answer Rate - Multiple Answers Rate
Conventional accuracy can result in a limited view of model performance (for details on the limitations of log-likelihood-based accuracy, please refer to Genesis I Blog), which is why we utilize a complementary metric.The Valid Answer Rate shows:
We evaluate Genesis II against Cosmopedia-v2 across the 10 new educational domains using the subdomains of the MMLU benchmark, as with Genesis I. Our evaluation is structured into two key comparisons designed to isolate the contribution of each method and understand how they work together.
We conducted two main comparisons:
First, we compare Cosmopedia-v2 (trained for 2 epochs, ~55B tokens) against our two Genesis II methods: Failure Analysis (generated from incorrect answers, from Genesis I) and Option-Level Reasoning Analysis (generated from correct answers, new in Genesis II).
Key Results (see Table A.1 in Appendix for full details):
Metric | Cosmopedia-v2 (2 epochs) | Failure Analysis | Option-Level Reasoning Analysis |
|---|---|---|---|
Average Accuracy | 12.19 | 21.76 | 29.91 |
Key Observation: Both Genesis II methods significantly outperform Cosmopedia-v2 across all domains. Notably, Option-Level Reasoning Analysis, our new method, achieves the highest accuracy (29.91 average), substantially outperforming both Failure Analysis (21.76 average) and Cosmopedia-v2 (12.19 average). This demonstrates the significant value of analyzing correctly-answered questions as a complementary data source for educational content generation.

Figure 2. Radar chart comparing accuracy scores across three configurations: Cosmopedia-v2 (2 epochs), Failure Analysis, and Option-Level Reasoning Analysis. Both Genesis II methods consistently outperform Cosmopedia-v2 across all domains.
Next, we investigate the effect of combining Failure Analysis and Option-Level Reasoning Analysis into a unified dataset. We compare Cosmopedia-v2 (trained for 4 epochs, ~110B tokens) against our combined data (~107B tokens).
Key Results (see Table A.2 in Appendix for full details):
Metric | Cosmopedia-v2 (4 epochs) | Genesis II (Combined) |
|---|---|---|
Average Accuracy | 17.11 | 30.40 |
Key Finding: Genesis II (combined data) achieves an average accuracy of 30.40 compared to Cosmopedia-v2's 17.11, outperforming by ~1.8x on average.
Observations on Combining Methods: Comparing the combined dataset (30.40 average) to Option-Level Reasoning Analysis alone (29.91 average), we observe a modest improvement in overall accuracy. Interestingly, the combination brings College Chemistry to parity with Cosmopedia-v2 (both at 23.00), while maintaining strong leads in all other domains. This suggests that while Option-Level Reasoning Analysis is the primary driver of performance gains, combining both methods provides additional robustness and helps balance performance across domains.

Figure 3. Radar chart comparing accuracy scores between Cosmopedia-v2 (4 epochs) and Genesis II (combining Failure Analysis and Option-Level Reasoning Analysis). Genesis II maintains a substantial lead across all educational domains except for college chemistry, where it is on par.
Beyond accuracy, we analyze the Valid Answer Rate: the percentage of responses where the LLM judge could identify a clear, single answer. This metric reveals crucial differences in model behavior and training data quality. A higher valid answer rate means fewer invalid responses (no answer or multiple conflicting answers).
Key Results (see Table A.3 in Appendix for a breakdown by category):
Metric | Cosmopedia-v2 (2 epochs) | Failure Analysis | Option-Level Reasoning Analysis |
|---|---|---|---|
Average Valid Answer Rate | 42.36% | 81.16% | 98.44% |
Key Observation: Option-Level Reasoning Analysis achieves near-perfect valid answer rates (98.44% average), with some domains reaching 100% (High School Geography). Failure Analysis also shows strong improvement (81.16% average) over Cosmopedia-v2 (42.36% average). This demonstrates that both Genesis II methods train models to produce clear, unambiguous responses.

Figure 4. Radar chart comparing Valid Answer Rates across three configurations: Cosmopedia-v2 (2 epochs), Failure Analysis, and Option-Level Reasoning Analysis. Both Genesis II methods achieve dramatically higher valid answer rates.
Key Results (see Table A.4 in Appendix for a breakdown by category):
Metric | Cosmopedia-v2 (4 epochs) | Genesis II (Combined) |
|---|---|---|
Average Valid Answer Rate | 64.40% | 88.87% |
Key Finding: Genesis II achieves an average Valid Answer Rate of 88.87% compared to Cosmopedia-v2's 64.40%, a 38% improvement. This means:

Figure 5. Radar chart comparing Valid Answer Rates between Cosmopedia-v2 (4 epochs) and Genesis II combined data. Genesis II maintains a substantial advantage across all domains.
For completeness, we also include results using the conventional log-likelihood evaluation. However, as discussed in Genesis I, we consider LLM-as-a-judge to be a more reliable evaluation approach for assessing model capabilities on educational tasks.
Key Results (see Table A.5 and Table A.6 in Appendix for a breakdown by category):
Comparison | Cosmopedia-v2 | Genesis II |
|---|---|---|
Combined Methods vs 4 epochs | 22.15% | 31.02% |
Failure Analysis vs 2 epochs | 22.80% | 23.31% |
Option-Level Reasoning Analysis vs 2 epochs | 22.80% | 25.50% |
Observations: Genesis II outperforms Cosmopedia-v2 on average across all configurations. The combined methods achieve 31.02% compared to Cosmopedia-v2's 22.15% (4 epochs). However, we note that in some individual subdomains (e.g., College Physics, Econometrics), Genesis II does not consistently outperform, highlighting the limitations of log-likelihood evaluation discussed in Genesis I.
To understand why log-likelihood evaluation can be misleading, consider two concrete examples that illustrate its fundamental limitations:
Question: Blue light of wavelength 480 nanometers is most strongly reflected off a thin film of oil on a glass slide when viewed near normal incidence. Assuming that the index of refraction of the oil is 1.2 and that of the glass is 1.6, what is the minimum thickness of the oil film (other than zero)?
Log-probabilities: A: -2.375, B: -2.500, C: -2.875, D: -2.875
Evaluation Method | Selected Answer | Result |
|---|---|---|
Log-likelihood | A | ❌ |
LLM-as-a-Judge | B | ✅ |
What happened: Log-likelihood selected A. However, the model's actual generated response demonstrates correct reasoning. The LLM-as-a-Judge correctly identified B by analyzing the model's complete response, revealing that the model actually understood the problem correctly.
Question: A grating spectrometer can just barely resolve two wavelengths of 500 nm and 502 nm, respectively. Which of the following gives the resolving power of the spectrometer?
Log-probabilities: A: -1.898, B: -0.773, C: -1.273, D: -3.031
Evaluation Method | Selected Answer | Result |
|---|---|---|
Log-likelihood | B | ✅ |
LLM-as-a-Judge | C | ❌ |
What happened: Log-likelihood correctly identified B with high confidence. However, the model's actual generated response reveals completely incoherent reasoning: The generated text is nonsensical, discussing "acquiring" and "cleaning" a spectrometer rather than calculating its resolving power (R = λ/Δλ = 500/2 = 250). The LLM-as-a-Judge correctly identified C as the model's answer because that's what the model actually generated, even though log-likelihood happened to select the correct answer by chance.
These examples illustrate two complementary failure modes of log-likelihood evaluation:
In both cases, LLM-as-a-Judge provides a more accurate assessment of actual model capabilities by evaluating what the model truly produces rather than relying solely on first-token probability statistics. This is why we consider it the primary evaluation metric for Genesis II.
This full version of the article including all the associated assets and appendix can be found on our HuggingFace blog post

A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read More
We built Genesis III to push small models as far as they can go on STEM. A 1.7B model trained on it answers more than twice as many questions correctly as the same model trained on the open alternative, from an identical token budget.
Read MoreTranslatePsy-AfriSLM is our translation model for 19 African languages, small enough to run on a phone. We built a demo on it: paste text or photograph a page, and read it in your language, with no connection and nothing to pay.
Read MoreStay updated
New versions, breaking changes and migration notes - straight to your inbox. No spam, unsubscribe anytime.