
Local AI 101: what your own machine can run
A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read MoreWe present a unified, cross-platform framework that successfully enables parameter-efficient training of modern LLMs with LoRA on consumer hardware such as mobile SoCs and desktop GPUs, without relying on a CUDA-only ecosystem.

The ability to fine-tune Large Language Models (LLMs) on user-specific data is crucial for personalization and broader adoption. To preserve privacy, ensure operational continuity in high-latency regions (e.g., emerging markets), and provide an anti-fragile, highly resilient, and scalable AI platform, it is desirable that this fine-tuning occurs locally on consumer devices. However, existing on-device fine-tuning solutions are limited: they lack GPU acceleration or are restricted to specific vendor ecosystems, failing to support the diverse range of consumer-grade and mobile hardware.
Tether Data, S.A. de C.V. (Tether Data, we, us, our) presents a portable, LoRA-based fine-tuning solution that operates across the full spectrum of consumer GPU architectures. By integrating fine-tuning directly into a cross-platform inference engine and leveraging a portable graphics API, our solution enables efficient training on diverse hardware, from smartphones (with Mali, Adreno, and Apple GPUs) to desktops, laptops and servers (with AMD, Intel, NVIDIA, and Apple GPUs). We also implemented the necessary architectural extensions and operator designs to support LoRA LLM fine-tuning for modern transformer models like Qwen3 and Gemma3.
We validate our approach with two real-world applications: email style transfer and biomedical question answering. The results demonstrate successful on-device fine-tuning across all tested platforms. By transforming LLM fine-tuning from a vendor-specific capability into a cross-platform solution, our work represents a significant step toward democratizing LLM personalization and making this technology accessible to a wide audience of end users.
The growing capabilities of Large Language Models (LLMs) have intensified the need for efficient fine-tuning methods to adapt these general-purpose models to specific tasks and domains. While full fine-tuning (which updates all of a model's parameters) offers maximum flexibility, its prohibitive computational and memory costs render it impractical for most users.
This challenge has accelerated the development of Parameter-Efficient Fine-Tuning (PEFT) methods, which achieve strong performance by updating only a small fraction of a model's weights. Among PEFT techniques, Low-Rank Adaptation (LoRA) has emerged as a dominant standard. By freezing a model's pre-trained weights and injecting trainable, low-rank matrices into its transformer layers, LoRA reduces the number of trainable parameters by orders of magnitude, making fine-tuning feasible on consumer-grade hardware. However, a significant accessibility gap remains: LoRA’s practical implementation is restricted to a limited set of GPU architectures, preventing broader use of the parallel processing power available in modern mobile GPUs. Most open-source finetuning frameworks, such as PEFT, Unsloth, and Llama factory, are built upon PyTorch or JAX and primarily target NVIDIA's CUDA platform. While highly effective, this focus perpetuates hardware dependency.
The llama.cpp project has become the de-facto library for efficient, cross-platform LLM inference, supporting a wide range of hardware from Windows, macOS, and Linux to mobile devices. Nevertheless, the existing implementation is constrained by several shortcomings:
Our work addresses the logical next step: democratizing LoRA fine-tuning across this diverse hardware ecosystem. We present the integration of LoRA directly into the llama.cpp framework. To achieve true cross-platform compatibility, we leverage the Vulkan [3] graphics and compute API, which provides a unified programming interface across virtually all modern GPU vendors, including desktop (AMD, Intel, NVIDIA) and mobile (Qualcomm Adreno, ARM Mali, Apple) platforms. Our work makes the following key contributions:
Through use cases in email style transfer and biomedical Q&A, we demonstrate our solution's core capabilities:
To enable the community to build upon this work, we are publicly releasing the following resources:
This work democratizes LLM fine-tuning by breaking its dependency on specific hardware vendors, offering a cross-platform solution for practitioners, developers, and end-users alike.
Architecture Overview

Figure 1 : We integrated the Low-Rank Adaptation (LoRA) module in llama.cpp augments the original pretrained weight matrix W with a low-rank update, scaled by alpha. This is expressed as:
W' = (W + AB) * alpha/r
Key Features
The LoRA adapters are applied to all linear layers within the transformer blocks. This includes the query, key, value, and output projections in the self-attention mechanism, as well as the linear layers in the feed-forward network (FFN). This comprehensive application was chosen to maximize the adaptive capacity of the finetuning process. To manage this process, we introduced a set of functions into the llama.cpp public API:

Figure 2: How inputs are transformed through LoRA adapters, added to the frozen base model weights W, and then forwarded as outputs.
Our choice of the Vulkan backend was strategic for achieving true cross-platform support. Unlike proprietary APIs like CUDA or Metal, Vulkan is a modern, low-level, vendor-agnostic standard that provides direct control over GPU hardware across a vast ecosystem. This includes all major desktop (NVIDIA, AMD, Intel) and mobile (Qualcomm Adreno, ARM Mali) vendors, making it the ideal choice to democratize fine-tuning. To leverage this backend for robust training, several foundational capabilities had to be added to llama.cpp:
New GGML Operations
GGML_OP_CROSS_ENTROPY_LOSS_MASKED | Computes cross-entropy loss only on unmasked tokens (assistant responses). |
GGML_OP_CROSS_ENTROPY_LOSS_MASKED_BACK | Performs the backward pass for masked loss, propagating gradients only where the mask is active. |
GGML_OP_COUNT_EQUAL_MASKED | Counts correct predictions only over masked (unmasked) tokens to compute meaningful accuracy metrics. |
All operations are compatible with GPU backends, including Vulkan, and have dedicated shader implementations for efficient masked reductions and elementwise operations.
These shaders are optimized for coalesced memory access and utilize fused multiply-add (FMA) instructions for high throughput.
Shader | Description |
|---|---|
cross_entropy_loss_masked_back.comp | Computes masked cross-entropy gradients efficiently on GPU, handling both masked and unmasked tokens. |
count_equal_masked.comp | Counts correct predictions in unmasked positions to compute accuracy metrics directly on GPU. |
New GGML API Functions
GGML_API struct ggml_tensor * ggml_cross_entropy_loss_masked(..);
GGML_API struct ggml_tensor * ggml_count_equal_masked(..);
GGML_API struct ggml_tensor * ggml_opt_dataset_masks(..);These new APIs expose masked operations and dataset functionality to downstream modules such as finetune-lora.cpp.
As part of the effort to support instruction fine-tuning capabilities in llama.cpp, we implemented masked-loss training, where a mask is applied to train only on assistant tokens. This enables the model to focus exclusively on assistant responses while ignoring system and user prompts - a critical component for instruction-following model alignment.
Masked Loss Computation
The centerpiece of the instruction fine-tuning implementation is masked loss, which ensures that only assistant responses contribute to the training objective. This is essential for instruction fine-tuning, where the goal is to optimize the model’s behavior as the assistant, not as the user or system. During dataset processing, each token in the sequence is annotated with a mask value indicating whether it should contribute to the loss function. Tokens corresponding to assistant messages are marked with 1, while all others (system or user tokens) are set to 0.
This design ensures that:
Chat Template System
Instruction fine-tuning depends on consistent conversational formatting. The implementation includes a chat template system supporting both the ChatML format and custom Jinja templates, ensuring compatibility with Hugging Face datasets and tokenizer pipelines. The ChatML format provides a structured markup separating roles (system, user, assistant) and ensures interoperability with ChatML-compatible datasets.
This flexible system allows training across diverse datasets without modifying the core pipeline.
The dataset pipeline handles preprocessing from JSONL inputs to fully tokenized, masked and padded tensors. It ensures consistent formatting and masking across various dataset sources.
Input Format Example
{
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of France?"},
{"role": "assistant", "content": "The capital of France is Paris."}
]
}Usage Examples
Basic Instruction Fine-tuning
# Using built-in ChatML template
./build/bin/llama-finetune-lora \
-m model.gguf \
-f conversations.jsonl \
--assistant-loss-only \
--lora-rank 16 \
--lora-alpha 32 \
-c 512 -b 128 -ub 128 \
-ngl 999 -fa offCustom Chat Template
# Using custom Jinja template
./build/bin/llama-finetune-lora \
-m model.gguf \
-f conversations.jsonl \
--assistant-loss-only \
--chat-template custom_template.jinja \
--lora-modules "attn_q,attn_k,attn_v,attn_o" \
-c 128 -b 128 -ub 128 \
-ngl 999 -fa offThese commands demonstrate how to enable masked-loss instruction fine-tuning with LoRA adapters, using either the built-in ChatML template or a custom dataset template.
A primary challenge was enabling LoRA finetuning for models like Qwen3 on Qualcomm Adreno GPUs (e.g., Adreno 830) using Vulkan. The investigation focused on crashes in the Vulkan backend during MUL_MAT and OUT_PROD operations involving very large tensors for LoRA fine-tuning. The team initially simplified the shaders to create a minimal baseline, then developed minimal test cases in the llama.cpp suite to reproduce the crashes.
After confirming the issue was linked to complex indexing arithmetic on large buffers within a minimal GLSL shader and noting that the OpenCL backend passed the same tests, the problem was isolated to the Vulkan driver. The root cause was finally identified as an undocumented limit in the Adreno 830 Vulkan driver concerning the cumulative size of input and output buffers (SSBOs) for a single operator.
We implemented a dynamic tiling algorithm for MUL_MAT and OUT_PROD. Instead of performing one large matrix multiplication, the operation is broken down into smaller, independent tiles that respect the 128MiB memory limit.

Figure 3: Dynamic tiling solution for very large matrices.
The algorithm is as follows:
This approach allows llama.cpp to execute arbitrarily large matrix operations on the Adreno GPU without triggering the hardware limitation, with the tile sizes adapting dynamically to different models and data types.
The primary result of this work is the successful enablement of LoRA finetuning for modern LLMs (Qwen3, Gemma) on cross-platform GPUs. The fine-tuned models are made available on the same license terms as the original model to facilitate a review of our work, software, and support research: Gemma models on the Gemma Terms of Use and Qwen models on the Apache 2.0 license.
We validated our work across a range of hardware and models on multiple datasets to evaluate unstructured finetuning and instruction finetuning:
We evaluate cross-backend LoRA fine-tuning with two complementary corpora chosen to stress different aspects of the training stack while remaining lightweight enough for mobile/edge GPUs.
Both datasets are synthetically generated to limit the likelihood of the inclusion of PII. The synthetic email data set is made available under the CC-BY-NC 4.0 (Creative Commons Attribution–Non Commercial 4.0), which is license-compatible for research, and preprocessed into compact JSONL. The structured, biomedical yes/no questions were specific entries in the PubMedQA dataset as described herein. The PubMedQA dataset was made available under the MIT license at the time it was accessed. All splits are stratified where applicable (seed=42) and exported with a MANIFEST.json for counts and provenance.
This is a 100% machine-generated US-English corpus created in-session from constrained prompts; no scraping, no mailbox ingestion. Intended for style-transfer/formatting robustness rather than factual knowledge.
Content & Format: 200 emails with fields {id, subject, body}. Subjects include varied surface forms (e.g., Re:/Fwd:, lowercase, single-emoji); bodies are compact (<300 words), casual voice, occasional quoted reply lines (> …), and generic venues (“the park,” “the gym”) to avoid identifiers.
Safety & Licensing: Synthetically generated to limit the likelihood of real names, addresses, companies, or unique venues. Intended for research only. Made available under the CC-BY-NC 4.0 (Creative Commons Attribution–Non Commercial 4.0).
Sources & Access. PubMedQA labeled (297 samples, 18K tokens) subsets accessed programmatically via Hugging Face [2]. MIT-licensed per dataset card. Fields used: question, final_decision (Yes/No/Maybe), long_answer.
To validate core causal language modelling capabilities, we fine-tuned Qwen-1.7B on an email corpus using standard next-token prediction. This experiment validates:

Figure 4: Training loss on Qwen-1.7B - email data.
To validate that LoRA training inside llama.cpp behaves equivalently to established framework workflows (e.g., PyTorch + HuggingFace), we performed a small-scale instruction-tuning experiment on a biomedical Q&A dataset. The goal of this evaluation was not to maximize accuracy, but to demonstrate that:

Figure 5: Training loss on Qwen-1.7B - biomedical data.
In Table 1 below, we show the time per epoch (column 1) and total training time (column 2) for various hardware configurations.
Hardware | Time/Epoch | Full Training (8 epochs) |
|---|---|---|
RTX 4090 | 5.5 min | 45 min |
AMD 7900 XTX | 13 min | 1.7 hrs |
Intel Arc A770 | 20 min | 2.7 hrs |
Apple M3 Pro | 40 min | 5.3 hrs |
Adreno 830 | 1h 40min | 13 hrs |
Mali G715 | 7h 40min | 61 hrs |
Table 1: Running times for fine-tuning on different architectures.
📊 View complete benchmarks with detailed metrics across all platforms
Table 2 below shows the model quality evaluation in terms of Win Rate (which answer is judged as superior by a capable LLM judge?), accuracy (which model has a better biomedical knowledge?) and cosine similarity (compared to a reference LLM output, how similar are the outputs from PyTorch and QVAC fine-tuned LLMs?). The conclusion: near-parity quality with established frameworks, but works on 8x more hardware platforms.
Metric | QVAC-fabric-llm | PyTorch/HuggingFace |
|---|---|---|
LLM-as-Judge Win Rate | 45-48% | 52-55% |
Biomedical Accuracy | 79-94% | 78-86% |
Cosine Similarity | 0.82 | 0.77 |
Table 2: Quality comparison metrics for PyTorch and QVAC models.
Key Takeaways
The model showed consistent domain-adaptation behavior across all GPUs tested, from mobile (Mali, Adreno) to desktop (Intel, AMD, Apple) to datacenter-class NVIDIA GPUs. llama.cpp’s LoRA pipeline produces the same functional model adaptation patterns seen in PyTorch, even at small scale, validating correctness of:
Importantly, this biomedical task highlights the broader utility of portable fine-tuning: the ability to adapt models for high-stakes, knowledge-intensive environments such as healthcare, scientific research, and regulated enterprise applications, even on devices that traditionally have not been considered “training-capable”. By enabling consistent LoRA training across NVIDIA, AMD, Intel, Apple Silicon, and mobile GPUs, we make domain adaptation accessible in contexts where data privacy and locality are crucial. This means sensitive datasets never need to leave the user’s device or institution, supporting compliance-driven deployment models.
Future work will focus on advancing the framework's efficiency and model support through several key avenues. We plan to expand quantization support by integrating formats such as GPTQ-INT8 and Q5_K_M, which offer superior trade-offs between computational speed and model fidelity. Kernel optimizations will continue, with efforts to enhance cache locality in the OUT_PROD shader and tailor workgroup parameters for core operations on mobile GPUs. We will also pursue lower-overhead memory management by eliminating staging buffers and adopting bindless descriptors to minimize CPU contention. Finally, we will investigate advanced compiler-level optimizations, such as operator fusion on Adreno architectures, to further increase training throughput and hardware utilization.
We present a unified, cross-platform framework that successfully enables parameter-efficient training of modern LLMs with LoRA on consumer hardware such as mobile SoCs and desktop GPUs, without relying on a CUDA-only ecosystem. We leverage Vulkan for cross-vendor acceleration (Mali, Adreno, Intel, AMD, NVIDIA) and Metal for Apple platforms, enabling fine-tuning across heterogeneous devices with a consistent user-facing API and training interface. Our contributions include the development of critical GPU kernels and backward passes to support SOTA architectures like Qwen3 and Gemma3, and the introduction of a masked-loss objective for effective on-device instruction-tuning. Furthermore, we overcome the fundamental barrier of mobile fine-tuning through a novel dynamic tiling method that manages severe memory constraints. Collectively, these innovations break the long-standing hardware limitations, as demonstrated by the first successful fine-tuning on mobile GPUs and universal compatibility across desktop architectures. The results validated through our use cases confirm that high-quality, local and private fine-tuning is no longer confined to powerful data centers but is now a viable and accessible capability for the broad ecosystem of consumer-grade hardware, thereby enabling the way for a new generation of personalized, highly-resilient, anti-fragile, and privacy-preserving on-device AI applications.
This full version of the article including all the associated assets can be found on our HuggingFace blog post

A phone or an 8 GB laptop handles most everyday AI jobs, and video needs 32 GB. How much memory local AI needs, and the command that tells you what yours has.
Read More
We built Genesis III to push small models as far as they can go on STEM. A 1.7B model trained on it answers more than twice as many questions correctly as the same model trained on the open alternative, from an identical token budget.
Read MoreTranslatePsy-AfriSLM is our translation model for 19 African languages, small enough to run on a phone. We built a demo on it: paste text or photograph a page, and read it in your language, with no connection and nothing to pay.
Read MoreStay updated
New versions, breaking changes and migration notes - straight to your inbox. No spam, unsubscribe anytime.