Find the right team.
Turn a customer message into a route.
fahrenheitMeet Zircon v2. Give it a situation and a set of options.
Get a calibrated probability for each. Right on your Mac.



Fluent isn’t the same as grounded.
That’s where Zircon comes in.

Zircon v2 is a 0.6B-parameter decision model that runs locally on Apple silicon via MLX. Choose between options, ask yes or no, or score on a scale. Version 2 adds math verification and conversation matching.
Route a request. Verify the math. Match a reply.
Focused decisions, without a round trip to the cloud.
Reported median response times: 30–45 ms per decision on a MacBook Pro (Apple M5). Speed varies by hardware.
ILLUSTRATIVE TIMING · 45 ms SHOWN AT 20× SLOWER SPEED
Turn a customer message into a route.
Check a worked calculation.
Match a response to its request.
Zircon v2 · Fahrenheit Research internal testing, September 2026. Results in your own use may differ.
| Specification | Value |
|---|---|
| Version | Zircon v2 · 0.6B · 8-bit |
| Released | September 2026 |
| Model size | About 634 MB |
| Runs on | Apple silicon · fully local |
| Output | A calibrated probability for every option |
| Response time | 30–45 ms per decision¹ |
| Decision types | Pick one, yes or no, score on a scale |
| Decision | Accuracy |
|---|---|
| What to do: reply, escalate, archive, or mark as spam | 93% |
| Notify now or hold | 99.7% |
| Priority on a 4-level scale | 94% |
| Pending steps: execute, hold, defer, or escalate | 100% |
| Metric | Zircon v2 |
|---|---|
| Accuracy | 80.6% |
| Brier score · lower is better | 0.043 |
¹ MacBook Pro (Apple M5), median. Speed varies by hardware.
² Fahrenheit Research internal testing, September 2026, on held-out emails.
³ Fahrenheit Research internal testing, September 2026, 2,000 decisions.
Model weights, benchmarks and setup instructions on Hugging Face.
ZIRCON v2 · 0.6B · MLX 8-BITBuilt for specific requirements from our enterprise and business clients.
Each model focuses on a particular domain, workflow or task.

1.7B parameters · 4-bit
Manufacturing knowledge. On your Mac.

1.7B parameters · 4-bit
Legal work, with a local first pass.

9.2B parameters · 4-bit
Marketing focus. From search to lifecycle.

67 million parameters
Know the document. Find its next step.
A thin model scopes the task first. A frontier model steps in only when needed.
Focused tools for different kinds of work. Explore each card for specifications, intended uses and evaluation context.
At Fahrenheit Research, we build models with a clear remit: specialize in the task, or check the answer. Our research spans Thin Language Models, grounding verification, and the boundaries of what AI can reliably check.
Thin Language Models built around a domain, designed for local deployment.
DEPTH OVER BREADTH.Verification models trained to check grounding, citations, and constraints.
A HEALTHY DOSE OF SKEPTICISM.Your data. Your hardware. A custom model and verifier, built with the lab.
BUILD WITH US ↗Tell us about your domain, your data, and where the model needs to run.
Zircon v2 is a 0.6B-parameter decision model from Fahrenheit Research, released in September 2026. Give it a situation and a set of options, and it returns a calibrated probability for every option. It is built for focused decisions inside a workflow, with inference running fully on-device on Apple silicon.
Zircon scores the options you supply rather than writing a free-form answer. You can use it to pick one option, make a yes-or-no decision, or score on a scale. For example, a language model can draft a reply while Zircon helps assess which action should follow. Version 2 also supports math verification and conversation matching.
Email and message triage is one tested use case: choose whether to reply, escalate, archive or mark a message as spam; decide whether to notify now or hold; assign a priority; or assess a pending step. The website animations show illustrative workflows and probabilities, not a live connection to the model.
The reported median response time is 30–45 milliseconds per decision on a MacBook Pro with Apple M5. This is model response time on the tested hardware, not a guarantee for every request or the total time your application takes. Hardware and the surrounding workflow affect the experience. The website slows the timing animation down so you can follow it.
In Fahrenheit Research’s September 2026 internal testing, Zircon v2 scored 80.6% accuracy and a 0.043 Brier score across 2,000 typed decisions. A lower Brier score is better: it measures how closely predicted probabilities match outcomes. On held-out emails, action selection scored 93%, notify-or-hold 99.7%, four-level priority 94%, and pending-step selection 100%. These are test results, not guarantees for new messages. Evaluate it on examples from your own workflow.
The Zircon v2 download is about 634 MB, with 8-bit weights, and runs on Apple silicon via MLX. The download size is not a statement of total runtime memory use. Start with the setup and usage information in the Zircon v2 Hugging Face repository ↗.
Zircon’s model inference runs locally on your Apple silicon device; it does not need a cloud round trip to score an option. Your application may still use external services for storage, logging or other models. Whether data leaves the device depends on how you build that wider workflow.
You choose how to use the probabilities. They can help rank options, route a request, or send an uncertain case for review. A high score does not guarantee a correct decision. Set review thresholds using your own evaluation data, especially before connecting the output to actions that send messages, spend money or change important records.
These models were built for specific requirements from our enterprise and business clients. Corundum focuses on manufacturing, Touchstone on legal work, Citrine on growth and marketing, and Mica on document classification, sensitive-text flags and routing. Each has its own runtime, intended uses and evaluation context. Explore the model cards for specifications and limitations.
A Lens is a workflow pattern: a small, specialized model reads and scopes a task first, handles what it can locally, and passes the harder parts to a frontier model when needed. It is a way to organize the work across models. Any external model calls—and the data sent with them—depend on your integration.
Yes. Start with the task you need to solve, representative examples, the hardware it must run on, and how you will judge a useful result. Tell us about your domain and data constraints at research@f-r.co. We can discuss whether a focused model fits your requirements.