QVAC Psy
QVAC Psy: The Foundation of Stable Intelligence
QVAC Psy is Tether’s family of state-of-the-art foundational models rooted in the principles of Psychohistory. Designed to provide a stable, objective substrate for a decentralized, composable, and infinitely scalable intelligence, QVAC Psy aims to ensure that intelligence remains a constant, unbreakable utility for all without a central authority.
import { loadModel, completion } from "@qvac/sdk"
const modelId = await loadModel({
modelSrc: "hf://qvac/MedPsy-1.7B-GGUF",
modelType: "llamacpp-completion"
})
const run = completion({
modelId,
history: [{ role: "user",
content: "Summarize these symptoms…" }],
stream: true
})Medical
QVAC MedPsy
QVAC MedPsy is the first specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC Psy medical models deliver reasoning capabilities previously exclusive to models seven times their size, setting a new benchmark for efficient, local intelligence.
Unprecedented efficiency
Across 7 medical benchmarks, our 1.7B model outperforms Google’s MedGemma 4B by over 11 points, despite being less than half its size. In rigorous testing on HealthBench Hard, the 1.7B model also outperformed the nearly sixteen-times-larger MedGemma 27B.
Sovereign power
On those same 7 medical benchmarks, our 4B model outperformed the nearly seven-times-larger MedGemma 27B.
Closing the parameter gap
This demonstrates that our smaller models, powered by a superior methodology, can match or outperform larger competing state-of-the-art models, achieving top-tier results on real-world medical and health clinical assessments.
0
medical benchmarks evaluated
0 pts
our 1.7B ahead of MedGemma 4B
~7×
smaller than the MedGemma 27B our 4B beats
Medical
Optimized for the edge
Beyond raw reasoning, we have achieved massive token efficiency. QVAC MedPsy-4B produces accurate answers using 3.2 times fewer tokens than backbone models. Alongside the full-precision models, we also release a full suite of efficient quantized variants.
Real-time performance
Lower token counts translate to lower latency and faster inference on edge devices.
Quantized models
Our 4-bit quantization retains accuracy within ~1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource-constrained edge devices.
Universal access
Bringing high-level intelligence to bandwidth-constrained and privacy-sensitive environments.
Vision
QVAC VisionPsy
QVAC VisionPsy is the second specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC Psy vision-language models bring multimodal understanding that used to require the cloud onto the phone in your pocket, at just 460M parameters.
Best-in-class vision
Across 17 multimodal benchmarks, our 460M model leads its weight class on 16, with the highest overall score of any ~0.5B vision model tested.
Punching above its weight
On ScienceQA, MM-IFEval and POPE, our 460M model outperforms every 0.75B to 1B model tested, including models 2.3 times its size.
Complete capability coverage
It leads all four capability areas measured: document understanding and OCR, visual perception, reasoning and knowledge, and instruction following and reliability, with its widest margins in reasoning and visual perception.
0M
parameters, in Nano and Flash variants
0 of 17
multimodal benchmarks led in its weight class
0.0×
the size of models it still outperforms
Vision
Built for real-time vision
Beyond raw accuracy, we have achieved a dramatic reduction in latency. QVAC VisionPsy-Nano-460M-Flash reads an image using as few as 64 visual tokens where comparable models use up to 1088, while retaining ~99% of the full model’s quality.
Real-time performance
Flash reaches the first token in 0.3 seconds on an iPhone 15 and 2.6 seconds on a Galaxy S25 Ultra, up to 25 times faster than comparable full-token models on the same device.
Lower memory footprint
Peak memory drops by 24 to 39% on Android and 41 to 65% on iPhone 15 against full-token baselines, freeing the headroom a phone actually has.
Quantized models
Alongside the full-precision models, we release GGUF quantized variants built for on-device inference through llama.cpp and the QVAC SDK.
Translation
QVAC TranslatePsy
QVAC TranslatePsy is the third specialized evolution of our foundational logic, purpose-built for machine translation. The family spans two models: TranslatePsy-AfriSLM for frontier quality on 19 African languages, and TranslatePsy-Nano for browser-scale footprint on 9 European and 8 African languages.
TranslatePsy-AfriSLM
Frontier quality on 19 African languages. It matches or beats systems up to 152 times its size.
TranslatePsy-Nano
Browser-scale footprint on 9 European and 8 African languages. It trades a little quality for reach, so translation runs on devices a large model never could.
0
African languages in AfriSLM
0×
larger systems AfriSLM matches or beats
0 + 8
European and African languages in Nano
Translation
Breaking the language barrier with AI
A language barrier is an access barrier. When medical knowledge, research or education is available only in a language you cannot understand, that knowledge may as well not exist. Translation removes that wall, which is why our model weights are open, for anyone to build on and extend to the languages we have not reached.
Get started
Run AI on what you own
One install, JavaScript or Python. 10+ AI tasks. Private, offline and free on the hardware you already own.