QVAC

QVAC Psy

QVAC Psy: The Foundation of Stable Intelligence

QVAC Psy is Tether’s family of state-of-the-art foundational models rooted in the principles of Psychohistory. Designed to provide a stable, objective substrate for a decentralized, composable, and infinitely scalable intelligence, QVAC Psy aims to ensure that intelligence remains a constant, unbreakable utility for all without a central authority.

medpsy.js
import { loadModel, completion } from "@qvac/sdk"

const modelId = await loadModel({
  modelSrc: "hf://qvac/MedPsy-1.7B-GGUF",
  modelType: "llamacpp-completion"
})

const run = completion({
  modelId,
  history: [{ role: "user",
    content: "Summarize these symptoms…" }],
  stream: true
})

Medical

QVAC MedPsy

QVAC MedPsy is the first specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC Psy medical models deliver reasoning capabilities previously exclusive to models seven times their size, setting a new benchmark for efficient, local intelligence.

  • Unprecedented efficiency

    Across 7 medical benchmarks, our 1.7B model outperforms Google’s MedGemma 4B by over 11 points, despite being less than half its size. In rigorous testing on HealthBench Hard, the 1.7B model also outperformed the nearly sixteen-times-larger MedGemma 27B.

  • Sovereign power

    On those same 7 medical benchmarks, our 4B model outperformed the nearly seven-times-larger MedGemma 27B.

  • Closing the parameter gap

    This demonstrates that our smaller models, powered by a superior methodology, can match or outperform larger competing state-of-the-art models, achieving top-tier results on real-world medical and health clinical assessments.

0

medical benchmarks evaluated

0 pts

our 1.7B ahead of MedGemma 4B

~7×

smaller than the MedGemma 27B our 4B beats

Medical

Optimized for the edge

Beyond raw reasoning, we have achieved massive token efficiency. QVAC MedPsy-4B produces accurate answers using 3.2 times fewer tokens than backbone models. Alongside the full-precision models, we also release a full suite of efficient quantized variants.

Real-time performance

Lower token counts translate to lower latency and faster inference on edge devices.

Quantized models

Our 4-bit quantization retains accuracy within ~1 point of BF16 while reducing disk footprint by ~69%, enabling practical deployment on resource-constrained edge devices.

Universal access

Bringing high-level intelligence to bandwidth-constrained and privacy-sensitive environments.

Vision

QVAC VisionPsy

QVAC VisionPsy is the second specialized evolution of our foundational logic, purpose-built for edge deployment. These QVAC Psy vision-language models bring multimodal understanding that used to require the cloud onto the phone in your pocket, at just 460M parameters.

  • Best-in-class vision

    Across 17 multimodal benchmarks, our 460M model leads its weight class on 16, with the highest overall score of any ~0.5B vision model tested.

  • Punching above its weight

    On ScienceQA, MM-IFEval and POPE, our 460M model outperforms every 0.75B to 1B model tested, including models 2.3 times its size.

  • Complete capability coverage

    It leads all four capability areas measured: document understanding and OCR, visual perception, reasoning and knowledge, and instruction following and reliability, with its widest margins in reasoning and visual perception.

0M

parameters, in Nano and Flash variants

0 of 17

multimodal benchmarks led in its weight class

0.0×

the size of models it still outperforms

Vision

Built for real-time vision

Beyond raw accuracy, we have achieved a dramatic reduction in latency. QVAC VisionPsy-Nano-460M-Flash reads an image using as few as 64 visual tokens where comparable models use up to 1088, while retaining ~99% of the full model’s quality.

Real-time performance

Flash reaches the first token in 0.3 seconds on an iPhone 15 and 2.6 seconds on a Galaxy S25 Ultra, up to 25 times faster than comparable full-token models on the same device.

Lower memory footprint

Peak memory drops by 24 to 39% on Android and 41 to 65% on iPhone 15 against full-token baselines, freeing the headroom a phone actually has.

Quantized models

Alongside the full-precision models, we release GGUF quantized variants built for on-device inference through llama.cpp and the QVAC SDK.

Translation

QVAC TranslatePsy

QVAC TranslatePsy is the third specialized evolution of our foundational logic, purpose-built for machine translation. The family spans two models: TranslatePsy-AfriSLM for frontier quality on 19 African languages, and TranslatePsy-Nano for browser-scale footprint on 9 European and 8 African languages.

  • TranslatePsy-AfriSLM

    Frontier quality on 19 African languages. It matches or beats systems up to 152 times its size.

  • TranslatePsy-Nano

    Browser-scale footprint on 9 European and 8 African languages. It trades a little quality for reach, so translation runs on devices a large model never could.

0

African languages in AfriSLM

0×

larger systems AfriSLM matches or beats

0 + 8

European and African languages in Nano

Translation

Breaking the language barrier with AI

A language barrier is an access barrier. When medical knowledge, research or education is available only in a language you cannot understand, that knowledge may as well not exist. Translation removes that wall, which is why our model weights are open, for anyone to build on and extend to the languages we have not reached.

Get started

Run AI on what you own

One install, JavaScript or Python. 10+ AI tasks. Private, offline and free on the hardware you already own.