New: VisionPsy-Nano, a 460M vision model that outperforms models twice its size.

Download Model

QVAC Psy: The Foundation of Stable Intelligence

QVAC Psy, is Tether's family of state-of-the-art foundational models rooted in the principles of Psychohistory. Designed to provide a stable, objective substrate for a decentralized, composable, and infinitely scalable intelligence, QVAC Psy aims to ensure that intelligence remains a constant, unbreakable utility for all without a central authority.

QVAC VisionPsy

QVAC VisionPsy is the second specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC PSY vision-language models bring multimodal understanding that used to require the cloud onto the phone in your pocket, at just 460M parameters.

Best-in-Class Vision: Across 17 multimodal benchmarks, our 460M model leads its weight class on 16, with the highest overall score of any ~0.5B vision model tested.

Punching Above Its Weight: On ScienceQA, MM-IFEval and POPE, our 460M model outperforms every 0.75B to 1B model tested, including models 2.3 times its size.

Complete Capability Coverage: It leads all four capability areas measured: document understanding and OCR, visual perception, reasoning and knowledge, and instruction following and reliability, with its widest margins in reasoning and visual perception.

Built for Real-Time Vision

Beyond raw accuracy, we have achieved a dramatic reduction in latency. QVAC VisionPsy-Nano-460M-Flash reads an image using as few as 64 visual tokens where comparable models use up to 1088, while retaining ~99% of the full model's quality.

Real-Time Performance: Flash reaches the first token in 0.3 seconds on an iPhone 15 and 2.6 seconds on a Galaxy S25 Ultra, up to 25 times faster than comparable full-token models on the same device.

Lower Memory Footprint: Peak memory drops by 24-39% on Android and 41-65% on iPhone 15 against full-token baselines, freeing the headroom a phone actually has.

Quantized Models: Alongside the full-precision models, we release GGUF quantized variants built for on-device inference through llama.cpp and the QVAC SDK.

QVAC MedPsy

QVAC MedPsy is the first specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC PSY medical models deliver reasoning capabilities previously exclusive to models seven times their size, setting a new benchmark for efficient, local intelligence.

Unprecedented Efficiency: Across 7 medical benchmarks, our 1.7B model outperforms Google’s MedGemma 4B by over 11 points, despite being less than half its size. In rigorous testing on HealthBench Hard, the 1.7B model also outperformed the nearly sixteen-times-larger MedGemma 27B.

Sovereign Power: On those same 7 medical benchmarks, our 4B model outperformed the nearly seven-times-larger MedgGemma-27B.

Closing the Parameter Gap: This demonstrates that our smaller models, powered by a superior methodology, can match or outperform larger competing state-of-the-art models, achieving top-tier results on real-world medical and health clinical assessments.

Optimized for the Edge

Beyond raw reasoning, we have achieved massive token efficiency. QVAC MedPsy-4B produces accurate answers using 3.2 times fewer tokens than backbone models. Alongside the full-precision models, we also release a full suite of efficient quantized variants.

Real-Time Performance: Lower token counts translate to lower latency and faster inference on edge devices.

Quantized Models: Our 4‑bit quantization retains accuracy within ~1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource‑constrained edge devices.

Universal Access: Bringing high-level intelligence to bandwidth-constrained and privacy-sensitive environments.

Your Infrastructure,
Your Sovereignty

Every QVAC Psy model is fully open-source, allowing deployment on your own infrastructure.

Full Open Source: Audit, refine, and adapt the weights to your specific needs. All our models, medical and vision, are released in both full-precision checkpoints and complete GGUF-quantized variants.

Local Execution: Achieve near-native speeds on local GPUs using the QVAC Fabric.

Stability by Design: A sovereign engine of intelligence that ensures your data remains private in any environment.

FAQ

MedPsy

MedPsy is a family of compact, text-only medical and healthcare large language models developed by Tether Data’s AI Research for edge and on-device deployment. The family includes MedPsy-1.7B, MedPsy-4B, and GGUF quantized versions for local inference.

VisionPsy

VisionPsy-Nano is a family of compact vision-language models developed by Tether Data's AI Research for edge and on-device deployment. The family includes VisionPsy-Nano-460M tuned for quality, VisionPsy-Nano-460M-Flash tuned for latency, and GGUF quantized versions for local inference.