
Local AI 101: what your own machine can run
A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read MoreThere is a need for publicly available, large-scale synthetic datasets that are rigorously curated. Genesis I is our first effort in this direction.

Recent advances in large language model (LLM) pretraining have increased the focus on curating high-quality web-scale datasets. Synthetic data, generated to emulate real-world text distributions, has become critical for LLM development. Microsoft’s Phi model [1] demonstrated the value of large-scale synthetic datasets, generating billions of tokens for pretraining. HuggingFace subsequently developed Cosmopedia [2] to replicate Phi-1.5. Despite these efforts, existing synthetic datasets remain insufficiently refined for training state-of-the-art LLMs that can compete with leading closed-source/proprietary models.
Furthermore, generating high-quality pre-training synthetic datasets is resource-intensive and expensive. As a result, dataset creation has been limited to well-funded corporations and major research institutions, restricting broader participation in the AI research community.
There is a need for publicly available, large-scale synthetic datasets that are rigorously curated. Such datasets can lower the barrier to entry for academic institutions, small research labs, and public organizations, enabling wider experimentation and innovation. Moreover, synthetic data can be tailored to cover critical educational and scientific domains, supporting specialized training aligned with real-world learning objectives.
To address these challenges, Tether Data, S.A. de C.V. (Tether Data, we, us, our) introduces QVAC Genesis I, a large-scale multi-domain educational synthetic dataset designed to support open, high-quality LLM pretraining. Our contributions include:
Our methodology consists of a four-stage pipeline designed to generate high-quality synthetic educational content through systematic error analysis and correction. The approach leverages state-of-the-art language models to create domain-specific educational materials that address common misconceptions and learning gaps.

Figure 1. Diagram of the pipeline for generating synthetic data: Seeds Data are used as input for the Quality Filter, whose output becomes the input for the Scaling QA phase, in which 4 questions + options + target are generated for each seed. Each of these questions moves on to phase two (Model Answering), where a proposed solution to that question is generated using LLM. Finally, only proposed solutions that differ from the target (Compare to Gold Label) move on to the last phase (Failure Analysis), where an analysis of the incorrect answer and the correct solution to the question is generated in four different styles (educational textbook, web articles, qa, conversational dialogue).
Objective. From the seed pools, generate multiple-choice questions per domain/level (e.g., college_biology) and then use a small SOTA model to produce an answer. Only incorrect model answers are forwarded to the failure-analysis stage. Where a final text is generated in four different styles (educational textbook, question-answers, web articles, and conversation dialogue), in which the incorrect solution is first analysed and then the correct solution is given.
Our approach focuses on systematically generating large synthetic question–answer (QA) data from unstructured scientific text. We begin with domain-specific seed passages drawn from medical, biological, physical, and mathematical sciences, ensuring broad conceptual coverage across diverse knowledge areas. Using a scaled prompting strategy, a large-capacity language model is instructed to generate multiple-choice QA pairs inspired by the topics of each seed passage. Each pair consists of a question, four options, and one correct answer.
The prompting process is dynamically adjusted to produce different levels of conceptual complexity, ranging from high-school fundamentals to college-level analytical reasoning. By modifying the prompt design, the same framework can be extended to generate domain-specific data of varying difficulty, enabling rapid expansion of high-quality training material for any scientific discipline. The resulting synthetic QA corpus is employed as annealing data, helping to refine the model during late-stage pretraining or fine-tuning for task alignment. These data are employed to perform inference-time evaluation and failure analysis across existing language models. By analyzing the types of questions where models consistently underperform such as reasoning-intensive, multi-concept, or numerically grounded items, we can systematically identify the weaknesses of each model.
This methodology demonstrates how scalable prompting can be leveraged to create domain-balanced, complexity-controlled synthetic QA data that supports both model assessment and future pretraining efforts. It bridges the gap between raw scientific text and structured evaluation resources, helping to reveal capability gaps in large language models across critical scientific domains. For detailed information about the prompt used see Appendix (Prompt Templates).
Our answer generation and extraction approach focuses on systematically identifying and analyzing where state-of-the-art models fail, providing valuable insights into model limitations and creating targeted training data. We employ a sophisticated LLM-as-a-Judge framework to extract answers from model responses, enabling comprehensive analysis of model performance across different problem types and complexity levels.
Objective: The primary goal is to observe the output of state-of-the-art models and systematically extract question-answer pairs where they fail, creating a rich dataset of model weaknesses and misconceptions that can be used for targeted training and improvement.
Methodology: We use a three-stage process for answer generation and extraction:
This methodology enables us to systematically capture model failures across different domains, creating a comprehensive dataset of model weaknesses that can be used for targeted training and improvement. For detailed information about response categories, extraction processes, and the complete LLM-as-a-Judge framework, see Section 4.2, and for evaluation prompt see Appendix (Prompt Templates).
Our failure analysis approach focuses on creating high-quality educational content by systematically analyzing where state-of-the-art models fail and generating comprehensive explanations that not only provide correct answers but also analyze the reasoning behind model failures. This creates rich, pedagogically valuable content that addresses common misconceptions and learning gaps.
Objective: The primary goal is to create high-quality synthetic data in four different styles where not only the correct answer is provided, but also a thorough analysis of state-of-the-art model failures is included, creating comprehensive educational content that addresses misconceptions and learning gaps.
Methodology: We employ a systematic approach to failure analysis that generates synthetic educational content in four distinct styles:
All four styles are generated from MCQ, the model's wrong answer, and the correct label.
This methodology demonstrates how systematic failure analysis can be leveraged to create domain-balanced, pedagogically-rich synthetic data that supports both model assessment and educational content generation. For detailed information about the four-style content generation process and specific prompt templates, see Appendix (Prompt Templates).
Domain-Level Balance:
Error Distribution Strategy:
Format Standardization:
Quality Assurance Measures:
Tooling. We orchestrate the end-to-end pipeline using distilabel [5] running against a vLLM inference server (vLLM Team, 2024).
Pipeline Orchestration:
Model Architecture. We used the following open-source models in the various stages:
Flow
We pre-train a 1.7B-parameter transformer (Qwen3 family) initialized from scratch with BF16 mixed precision and context length 4,096. Tokenization uses the Qwen3 tokenizer; data are stored in HuggingFace Datasets (Arrow). The corpus totals 41B tokens (multi-domain) and is traversed for 1 epoch via a PyTorch DataLoader. To aid stability and throughput expected in technical deployments, we enable activation checkpointing, fused kernels where available (fused attention/optimizer), enable FlashAttention2 on H100, and torch.compile (safe mode) once the run is stable.
Optimization follows AdamW (weight decay 0.01), learning rate 2e-4, warmup 600 steps, gradient clipping 1.0, and seed 42. Per-GPU micro-batch is 4 with gradient accumulation 8 across 480 GPUs, yielding an effective global batch of 4×8×480=15,360 samples/step. We log train metrics every 50 steps, validate every 500 steps (20 eval iters), checkpoints are created every 1000 steps, and support resume with exact optimizer/state restoration. We achieved a total training throughput of 1.5 seconds per step (). We note common failure modes and mitigations: BF16 overflow (addressed via dynamic loss scaling), NCCL stalls (timeouts and interface pinning), and fragmentation (CUDA max_split_size_mb=512, expandable segments, GC threshold 0.8).
The resulting pre-trained model is publicly released and available at https://huggingface.co/qvac/genesisI-model
We made multiple training runs on 60 nodes with 8× NVIDIA H100 80GB per node (480 GPUs total), 8 CPUs per task, ~800 GB RAM per node, Slurm priority partition, exclusive allocation, and 72-hour time limit. We launch with srun using PyTorch DDP (world size 480), auto-detect the master from Slurm, and bind ranks to GPUs via Slurm’s environment. Stdout/stderr are streamed to logs_training/qvac_60node_training_%j.{out,err}; checkpoints are sharded and saved periodically for robust resume.
Networking is NCCL over InfiniBand with UCX transports. We use infiniband and set NCCL_IB_DISABLE=0 , NCCL_IB_HCA="mlx5", NCCL_SOCKET_IFNAME, and NCCL_BLOCKING_WAIT=1 with a 720-second watchdog to fail fast on fabric issues. UCX is configured for multi-device transport; we also pin file system threads and enable asynchronous I/O prefetch to keep GPUs fed.
Reliability & observability: W&B captures metrics, system traces, and artifacts; we additionally export structured logs (throughput, TFLOPs/GPU, GPU/host memory, step time. For reproducibility, we fix seeds, log exact launch scripts and env, and report effective tokens/step and utilization.
Volume, Diversity, and Domain Coverage:
Domain | Number of Samples | No of Tokens (in B) |
|---|---|---|
High school biology | 3,818,070 | 4.511 |
College biology | 3,286,648 | 3.927 |
Professional medicine | 1,552,474 | 1.884 |
College medicine | 5,164,247 | 6.218 |
High school mathematics | 3,244,240 | 4.277 |
College mathematics | 5,895,052 | 8.243 |
High school physics | 2,277,880 | 3.061 |
College physics | 4,281,062 | 5.814 |
Conceptual physics | 2,354,184 | 2.973 |
Total | 31,873,857 | 40.906 |

Figure 2. Histogram showing the results obtained using LLM-as-a-Judge method using Opencompass framework. The different educational domains of the MMLU dataset on the x-axis and the score on the y-axis. We can see that Qvac Genesis I performs better on average than the current largest synthetic dataset, Cosmopedia, and also in all individual topic and level domains except college physics.

Figure 3. Different representation of results obtained using LLM as a judge via the OpenCompass framework.
We developed a robust and stable evaluation framework using OpenCompass [8] that leverages LLM-as-a-Judge methodology to extract answers from model outputs. This approach represents a significant advancement over traditional log-likelihood-based evaluation methods commonly used in benchmarking.
Traditional Log-Likelihood Limitations:
Our LLM-as-a-Judge Approach: Our evaluation framework addresses these limitations by implementing a three-stage process:
This methodology provides several advantages:

Figure 4. Diagram of our evaluation pipeline. Stage 1: The model to be evaluated generates a complete response to the evaluation question. Stage 2: A specialized LLM judge extracts the final answer from the model's complete response. Stage 3: The extracted answer is compared against the ground truth using exact string matching. In the end, for each output we will have: Correct, Incorrect, Multiple Answer or No Answer.
Answer Extraction Framework
We implemented a sophisticated answer extraction system that can handle various response patterns and edge cases:
Response Categories:
Extraction Process: The LLM judge analyzes the complete model response to identify:
Our evaluation system uses a carefully designed prompt template that ensures consistent and reliable answer extraction. For detailed information about the complete prompt template, see Prompt Templates.
Primary Metrics:
Quality Assurance:
This evaluation framework provides a more accurate and comprehensive assessment of model performance, particularly for complex reasoning tasks where traditional log-likelihood methods may not capture the full extent of model capabilities.

Figure 5. Histogram showing the results obtained using the Loglikelihood method and LM-Harness framework. Even here we can see that Qvac Genesis I performs better on average than the current largest synthetic dataset, Cosmopedia, and also in all individual topic and level domains except college physics.

Figure 6. Different representation of results obtained using the Loglikelihood method and LM-Harness framework.
This full version of the article including all the associated assets and appendix can be found on our HuggingFace blog post

A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read More
We built Genesis III to push small models as far as they can go on STEM. A 1.7B model trained on it answers more than twice as many questions correctly as the same model trained on the open alternative, from an identical token budget.
Read MoreTranslatePsy-AfriSLM is our translation model for 19 African languages, small enough to run on a phone. We built a demo on it: paste text or photograph a page, and read it in your language, with no connection and nothing to pay.
Read MoreStay updated
New versions, breaking changes and migration notes - straight to your inbox. No spam, unsubscribe anytime.