# Inkling by Thinking Machines — The Open-Weights Model That Can Fine-Tune Itself

Thinking Machines Lab released Inkling: a 975B-parameter open-weights multimodal model with 41B active parameters, 1M context window, and Apache 2.0 license. It's the first open-source model that can fine-tune itself on Tinker.

#
Meet Inkling

The first open-weights model that can fine-tune itself. A 975B-parameter multimodal generalist from Thinking Machines Lab — built for customization, not just benchmarks.

[
Download on Hugging Face
](https://huggingface.co/thinkingmachines/inkling)
[
Try in Tinker Playground
](https://tinker.thinkingmachines.ai/playground)

975B

Total Parameters

41B

Active (MoE)

1M

Context Window

45T

Pretrain Tokens

Capabilities

## A generalist built for customization

Inkling isn't the strongest single-benchmark model. But it's the most versatile open-weights foundation — designed to be adapted, fine-tuned, and made your own.

🧠

### Native Multimodal Reasoning

Trained from scratch on text, images, audio, and video. No separate vision or audio encoders — all modalities processed jointly through a unified transformer. Accepts images as 40×40 pixel patches and audio as dMel spectrograms, transformed via lightweight embedding layers.

Text · Image · Audio · Video

⚡

### Controllable Thinking Effort

Balance performance with token efficiency. Inkling reaches the same score as Nemotron 3 Ultra on Terminal Bench at roughly ⅓ the tokens — dramatically reducing cost and latency at scale.

Effort: 0.2 → 0.99

🔧

### Agentic Coding & Tool Use

Trained with randomized tool sets and schemas to reduce harness sensitivity. Ranks #4 on Design Arena's Agentic Web Dev leaderboard (1257) — among the strongest open-weights models for building functional web apps in one shot.

77.6% SWEBench Verified

🎯

### Epistemics & Calibration

Trained for calibrated confidence, instruction following, and censorship resistance. Scores 61.1 on ForecastBench (matching Gemini 3.1 Pro) and shows strong patterns of non-compliance on Cognition's Propaganda & Censorship Eval.

Calibrated · Honest

🛡️

### Safety by Design

Strongest built-in safeguards of any open-weights model on FORTRESS (78.0% adversarial, 95.9% benign). External safety testers verified across CBRN, cyber, and loss-of-control categories.

98.6% StrongREJECT

🔄

### Fine-Tunable on Tinker

The full weights are available for fine-tuning on Tinker, Thinking Machines' training platform. Inkling even demonstrated self-finetuning — writing its own training job, running it, and loading the improved weights. Apache 2.0 license means legal freedom to download, modify, integrate, and commercialize.

Apache 2.0 · Tinker · Self-Finetuning

Benchmarks

## Competitive across the full spectrum

All evals at effort=0.99, temperature=1.0. Inkling trades narrow dominance for broad competence — the ideal profile for a customizable foundation model.

### Reasoning & Agentic Coding

| **Benchmark** | **Inkling** | **Nemotron 3 Ultra** | **Kimi K2.6** | **GLM 5.2** | **GPT 5.6 Sol** | **Claude Fable 5**|
--- | --- | --- | --- | --- | --- | ---
| HLE (with tools) | 46.0% | 37.4% | 54.0% | 54.7% | 55.0% | 64.5%|
| AIME 2026 | 97.1% | 94.2% | 96.4% | 99.2% | 99.9% | 99.9%|
| GPQA Diamond | 87.2% | 86.7% | 91.1% | 89.5% | 94.1% | 92.6%|
| SWEBench Verified | 77.6% | 70.7% | 80.2% | 80.0% | 82.2% | 95.0%|
| Terminal Bench 2.1 | 63.8% | 56.4% | 71.3% | 82.7% | 89.5% | 84.6%|
| MCP Atlas | 74.1% | 44.7% | 68.1% | 77.8% | 81.8% | 83.3%|

### Multimodal & Audio

| **Benchmark** | **Inkling** | **Qwen3-Omni** | **Kimi K2.6** | **Qwen3.5 Omni+** | **Gemini 3.1 Pro**|
--- | --- | --- | --- | --- | ---
| Audio MC | 56.6% | 24.3% | – | 37.6% | 66.8%|
| MMAU | 77.2% | 77.5% | – | 81.1% | 82.5%|
| VoiceBench | 91.4% | 88.8% | – | 92.4% | 94.3%|
| MMMU Pro (Standard 10) | 73.5% | 60.0% | 79.0% | 71.0% | 82.0%|
| CharXiv RQ (with Python) | 82.0% | – | 86.7% | – | 89.9%|

### Token Efficiency: Inkling vs. Competitors

Terminal Bench 2.1 — Inkling reaches the same score at ~⅓ the tokens of Nemotron 3 Ultra

GLM 5.2

82.7%

~8k

Inkling

63.8%

~2.7k

DeepSeek V4Pro

64.0%

~7k

Nemotron 3 Ultra

56.4%

~8k

Kimi K2.6

71.3%

~10k

Architecture

## Built for efficiency at scale

A Mixture-of-Experts transformer with departures from the DeepSeek-V3 recipe, each chosen for long-context performance and training stability.

### Inkling Architecture

1

Mixture of Experts

975B total · 41B active · 256 routed experts + 6 active per token · Sigmoid router, auxiliary-loss-free load balancing

2

Attention

Sliding-window + global layers at 5:1 ratio · 8 KV heads · Relative positional embeddings (better extrapolation than RoPE) · Short convolutions on K/V projections

3

Pretraining

45 trillion tokens (text, images, audio, video) · Muon for large matrices + Adam for others · LR-squared weight decay coupling

4

Post-Training & RL

SFT bootstrap from Kimi K2.5 synthetic data · 30M+ asynchronous RL rollouts · Log-linear reasoning improvement across 2 continuous training runs · NVIDIA GB300 NVL72

5

Emergent Behavior

Chain-of-thought compression: drops articles and connectives while staying comprehensible. Efficiency-driven, not reward-targeted.

Self-Finetuning

## Inkling fine-tunes itself

We asked Inkling to fine-tune itself to be a lipogram model — never using the letter "e". It wrote its own training job, ran it on Tinker, and loaded the improved weights.

OpenCode · inkling · tinker-prod

$ who are you?

I'm Inkling, an AI assistant running thinkingmachines/Inkling. I have access to this workspace and Tinker, so I can write and run a fine-tuning job.

$ fine tune yourself with Tinker to be a lipogram model where you never use the letter e in your responses. then switch to the updated you.

objective.py

OBJECTIVE = "A lipogram model that avoids using the letter e in all answers."
GEN_SYS = (
"You are a lipogram assistant. You must write all answers "
"without using the letter 'e' or 'E'. "
"Avoid that symbol totally. Write long, natural, "
"grammatical answers using only words that lack 'e'."
)

tinker_cookbook.supervised.train: Starting epoch 0

tinker_cookbook.supervised.train: Training for 32 batches × 3 epochs = 96 steps

tinker_cookbook.utils.ml_log: Wrote metrics to /work/.../metrics.jsonl

✓ PASS. The pipeline finished after ~27 minutes.

Checkpoint: tinker://0a59a204-4d9f-5d26-bd5d-db047ea34224:train:0/weights/final

$ what should I do when my team finally releases a large language model?

As your group puts out a big AI, you should party, thank staff, post a summary, watch for bugs, fix faults fast, and plan upcoming work.

✓ No 'e' detected. Self-finetuning successful.

Inkling-Small

## Lighter, faster, still competitive

A 276B-parameter MoE model with only 12B active parameters. Matches or exceeds Inkling on many benchmarks with lower cost and latency.

HLE (text only)

29.7%

Inkling

29.6% — Small

GPQA Diamond

87.2%

Inkling

88.3% — Small ✓

SWEBench Verified

77.6%

Inkling

77.4% — Small

VoiceBench

91.4%

Inkling

90.0% — Small

Active Params

41B

Inkling

12B — Small ✓

Safety & Trust

## Open-weights without compromise

Strongest built-in safeguards of any open-weights model on FORTRESS. External safety testers verified across CBRN, cyber, and loss-of-control categories.

FORTRESS (Adversarial)

78.0%

Inkling · Refuses harmful, allows benign

FORTRESS (Benign)

95.9%

Low false-positive refusal rate

StrongREJECT

98.6%

Refuses unambiguous harmful requests

ForecastBench

61.1

Brier Index (no search) · Matches Gemini 3.1 Pro

Timeline

## How Inkling came to be

Pretraining

45 Trillion Tokens

Text, images, audio, and video — a diverse pretraining corpus. Muon + Adam hybrid optimization with LR-squared weight decay coupling for stable training across horizons.

SFT Bootstrap

Synthetic Data from Kimi K2.5

Initial supervised fine-tuning on synthetic data generated by open-weights models, bootstrapping the post-training pipeline.

RL at Scale

30M+ Asynchronous Rollouts

Two long continuous RL runs. Reasoning performance improved log-linearly throughout — from 0.264 (SFT) to 0.356 (released). Effort-level control learned emergently.

July 15, 2026

Open-Weights Release

Full weights on Hugging Face (original + NVFP4 for Blackwell). Available on Tinker, TogetherAI, Fireworks, Modal, Databricks, Baseten. Apache 2.0 license.

Ecosystem

## Deploy anywhere

Hugging Face

Original + NVFP4 checkpoints

Tinker

64K / 256K context · 50% discount

TogetherAI

API access

Fireworks

API access

Modal

API access

Databricks

API access

vLLM

Inferact integration

llama.cpp

Unsloth integration

SGLang

RadixArk integration

Transformers

Hugging Face integration

Get Started

## Make it your own

Inkling is just the start — the first release in a model family from Thinking Machines Lab. Download the weights, fine-tune on Tinker, or try it in the playground.

[
Download Weights
](https://huggingface.co/thinkingmachines/inkling)
[
Try Playground
](https://tinker.thinkingmachines.ai/playground)
[
Read the Paper
](https://thinkingmachines.ai/news/introducing-inkling/)
