Is Inkling AI the Ultimate Open Source Model? Full Test
A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.
Ray Codes11 minTranscript found
Quick learning frame
Read this before watching.
Creative automation accelerates production while keeping human taste in brief, source selection, generation, editing, and critique.
New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Skill you build: The ability to design test prompts that probe an AI model across coding execution, calibrated forecasting, and long-form formatting, rather than relying on a single benchmark number.
Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.
Concept diagram
Where this video fits.
01Brief
02Source material
03Generation
04Selection
05Edit
06Taste review
07Reusable recipe
Deep lesson
Turn this video into working knowledge.
2,275 cleaned transcript words reviewed across 670 timed caption segments.
Thesis
Is Inkling AI the Ultimate Open Source Model? Full Test teaches a practical creative automation move: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.
The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.
0:00
Sparse but Massive
“Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale...”
Inkling is a 975B-parameter mixture-of-experts model that only activates 41B parameters per query (with a lighter 'Inkling small' variant using 12B active parameters), supports a 1M-token context, and reasons natively over text, image, and audio without converting audio to text first; it was even used to fine-tune itself entirely on its own, writing its own synthetic training data to learn a constrained 'no letter E' language task. Write down Inkling's active-vs-total parameter split (41B of 975B) next to a model you already use, to build intuition for how mixture-of-experts routing changes the hardware you actually need.
4:42
Agentic Coding Muscle
“functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt...”
Inkling was purpose-built for agentic coding and tool use, with official demos showing it build a full web app from a single prompt and then use an embedded browser agent to click through and test it, plus a multiplayer snake game with real-time server bots; it matches Nemotron 3 Ultra's score on the Terminal Bench 2.1 coding benchmark while using only a third of the tokens, and it's deployable locally via SGLang, VLLM, or Unsloth, with a special NVFP4 checkpoint for Blackwell GPUs and a 98% strong-reject safety score. Pick one of the three deployment paths (SGLang, VLLM, Unsloth) and read its docs for loading a model with an adjustable thinking-effort setting, since that's the lever Inkling exposes for trading cost against quality.
7:40
Calibrated Forecasting
“handle all the complex logic and the strict constraints that I had given in the initial prompt. And the second task was to check whether the human will land on Mars and return to Earth by 2035. And...”
Asked to forecast the probability of a human Mars round trip by 2035, Inkling didn't just guess: it gave roughly a 1% probability, walked through launch-window arithmetic, SpaceX's 'no earlier than 2028' Starship timeline, and life-support and certification requirements, then explicitly stated what evidence (like reusable orbital refueling by 2028) would raise its estimate above 10%. Give a reasoning model your own uncertain prediction question and require it to state both a probability and the specific evidence that would change that probability, then check whether its stated conditions are falsifiable.
01
Brief
Start with this video's job: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task. Treat "Brief" as the outcome you are trying to make visible, not a topic label. Anchor it to 0:00, where the video says: “Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale...”
02
Source material
Use "Source material" to locate the part of the creative automation mechanism the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 4:42, where the video says: “functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt...”
03
Generation
Turn "Generation" into the reusable artifact for this lesson: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints. This is where watching becomes something you can inspect and reuse.
04
Selection
Use "Selection" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.
05
Edit
Use "Edit" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.
06
Taste review
Use "Taste review" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.
07
Reusable recipe
Connect "Reusable recipe" to Is Inkling AI the Ultimate Open Source Model? Full Test by naming the claim, the evidence, and the artifact it should produce.
Example
Source-backed artifact packet
Convert the video into a scoped artifact request that includes the transcript claim, mechanism, acceptance criteria, and proof. The output should be a creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints..
Example
Creative automation proof brief
Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the creative automation pattern.
Example
Teach-back module
Transform the lesson into a definition, a Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe diagram, one misconception, one practice exercise, and a check-for-understanding question.
Do not learn it wrong
Treating the title as the lesson without checking what the transcript actually says.
mistaking novelty for quality
no source/brief discipline
shipping generated media without taste review
Letting the lesson drift into generic content advice.
Letting the lesson drift into tool hype.
Letting the lesson drift into creative output without selection criteria.
Do not count this as learned until these are true.
01
State the transcript-backed claim in your own words: A hands-on test of Inkling, a 975B-parameter (41B-active) mixture-of-experts model built for agentic coding and tool use, covering its self-fine-tuning demo, native multimodal reasoning, and three live tests: a Kanban app build, a calibrated Mars-2035 probability forecast, and a formatted long-form writing task.
02
Explain the practical stakes without hype: New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.
03
Map the idea onto the Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe sequence and name the weakest link.
04
Produce the artifact and include the evidence that proves it: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.
Put it into practice
Give this grounded prompt to Codex or Claude after watching.
You are helping me turn one specific YouTube video into real, durable learning.
Source video:
- Title: Is Inkling AI the Ultimate Open Source Model? Full Test
- URL: https://www.youtube.com/watch?v=U7IX2607jCM
- Topic: Creative Automation
- My current learning frame: Run the same three-test battery (a self-contained coding build, a calibrated probability forecast, and a long-form formatted writing task) against a local open-source model you can access, and compare its outputs to Inkling's results described in the video.
- Why this matters: New playlist item from Ray Codes; queued for transcript-backed review, topic mapping, and a practical learning artifact.
Transcript anchors from this exact video:
- 0:00 / Evidence 1: "Writing code with AI is great, but getting it to use external tools is much harder. Today we are looking at Inkling, a newly dropped model that excels at agentic workflows. Let us look at the sheer scale..."
- 2:52 / Evidence 2: "ultimate coding test where I'm going to test the agentic coding and the tool use capabilities of the Inkling model to see if it is actually able to build functional applications. I'm asking it to build a full..."
- 4:42 / Evidence 3: "functional web application in single short just from a simple text prompt. And then it used an embedded browser agent to actually click and interact with the final application. But it is not just good at single prompt..."
- 7:40 / Evidence 4: "handle all the complex logic and the strict constraints that I had given in the initial prompt. And the second task was to check whether the human will land on Mars and return to Earth by 2035. And..."
- 10:33 / Evidence 5: "raw multi-model power with the local hardware limits. I'll be putting the link to the official blog, the steps to install this model in your local system, or to try it on online platforms, and the three test..."
Video-aware target:
- Prompt lane: Creative automation
- Mechanism to extract: Extract the creative production loop, especially where the human keeps taste, selection, and final judgment.
- Artifact to produce: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.
- Artifact must include: brief; source inputs; generation recipe; selection criteria; edit/review checkpoint
Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, transcript support, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable mechanism from the video: Extract the creative production loop, especially where the human keeps taste, selection, and final judgment. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.
5. Include:
- a plain-English definition of the core idea
- a diagram or structured model using this sequence: Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe
- answers to these source questions: What asset is being produced? | What inputs and tools drive it? | Where does human taste intervene?
- 3 concrete examples that apply the video idea to real agentic work, such as Claude-generated video campaign; image-to-site workflow; voice or video editing loop
- 2 failure modes the video helps prevent, chosen from the transcript evidence and these likely risks: mistaking novelty for quality; no source/brief discipline; shipping generated media without taste review
- a checklist for the next real workflow, focused on: brief, inputs, generation, selection, critique
- one practical exercise with a clear done signal: Build one reusable creative recipe and define what would make the result rejectable.
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.
Quality bar:
- Make this specific to "Is Inkling AI the Ultimate Open Source Model? Full Test", not a generic Creative Automation essay.
- Anchor each creative step to transcript evidence about inputs, model/tool choices, iteration, editing, or critique.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- Avoid these generic drifts: generic content advice; tool hype; creative output without selection criteria.
- If evidence is weak or missing, stop and say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.
Misconceptions
What to stop believing.
Creative AI removes the need for taste.
It increases the need for taste because output volume explodes.
The best prompt is enough.
References, critique, iteration, and post-production matter just as much.
Practice studio
Learning only counts when you make something.
01
Transcript evidence map
Separate what the video actually says from what you already believe about the topic.
3 source-backed takeaways with timestamps, confidence, and a transfer note.02
One useful artifact
Apply the video to a real workflow and produce a creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints..
A reusable artifact with a done signal and one verification step.03
Creative automation teach-back card
Explain the creative automation mechanism to someone who has not watched the video yet.
A 90-second explanation, one diagram, one example, and one misconception to avoid.
Recall check
Answer first, then reveal — without rewatching.
How many of Inkling's 975 billion total parameters are actually activated per query, and why does that matter for running it on consumer hardware?
What benchmark result does the video cite to show Inkling's coding efficiency compared to Nemotron 3 Ultra?
What probability did Inkling assign to a human landing on Mars and returning safely by 2035, and what would change that estimate?
Source shelf
Use the video as a doorway, then verify with primary sources.