Creative Automation / Foundation

GLM 5.2 - The Top NEW Open Weights Model

Sam Witteveen examines GLM 5.2 after Z.AI released both full and FP8 open weights, walking through its Artificial Analysis results (beaten only by GPT 5.5, Opus 4.8, and the withdrawn Fable 5), its long chain-of-thought token habits, $1.40/$4.40 per-million pricing, and hands-on tests from the pelican SVG to a 5,000-word essay and a design-arena-worthy homepage.

Sam Witteveen13 minTranscript found

Quick learning frame

Read this before watching.

Creative automation accelerates production while keeping human taste in brief, source selection, generation, editing, and critique.

New playlist item from Sam Witteveen; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Skill you build: The ability to judge a newly released open-weights model beyond leaderboard rank — checking whether weights actually shipped, how token verbosity inflates results and costs, provider and data-residency options, and how it performs on your own hands-on tasks.

Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.

Concept diagram

Where this video fits.

01Brief
02Source material
03Generation
04Selection
05Edit
06Taste review
07Reusable recipe

Deep lesson

Turn this video into working knowledge.

2,360 cleaned transcript words reviewed across 656 timed caption segments.

Thesis

GLM 5.2 - The Top NEW Open Weights Model teaches a practical creative automation move: Sam Witteveen examines GLM 5.2 after Z.AI released both full and FP8 open weights, walking through its Artificial Analysis results (beaten only by GPT 5.5, Opus 4.8, and the withdrawn Fable 5), its long chain-of-thought token habits, $1.40/$4.40 per-million pricing, and hands-on tests from the pelican SVG to a 5,000-word essay and a design-arena-worthy homepage.

The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.

0:15

Weights actually released

“around with it this afternoon and the thing that just tipped me over the edge is artificial analysis have just published their stats on it as well. So one of the reasons I've been reluctant to make videos...”

Z.AI shipped both full and FP8 GLM 5.2 weights within 24 hours — notable because Chinese labs had been slow to release weights — though base models stay private, which Witteveen calls fair since paid base-model access (like Cursor fine-tuning Kimi through Fireworks) is how these companies profit; the model is post-trained for long-horizon tasks and adds multi-token prediction for speed. Before covering or adopting any 'open' model, verify what was actually released: instruct weights, FP8 variants, and whether the base model is available or held back commercially.

6:35

The token-count caveat

“intelligence with short amounts of tokens in here. All right, the last one that I thought was really kind of interesting, too, was the design arena. So, they've got actually put this model at the front above sort...”

GLM 5.2's Artificial Analysis jump over 5.1 puts it above DeepSeek, Qwen 3.7 Max, and MiniMax M3, but it earns that partly through very long chains of thought — outputting more tokens than DeepSeek, Qwen, and even Fable — while the industry (notably OpenAI since GPT 5.1) is heading the opposite way: high intelligence with fewer tokens; he also notes Fable 5 benchmarked without its 4.8 fallback does poorly due to nonsense rejections. When comparing models on an intelligence index, plot score against output tokens and ask whether the 'smarter' model is just buying accuracy with verbosity you'll pay for.

12:14

Hands-on and pricing math

“check out. This could be an alternative to a lot of the proprietary models that you're using. I really don't see the point in using some of the models, perhaps like a Sonnet, perhaps like a Gemini Flash,...”

At a uniform $1.40 in / $4.40 out per million tokens across providers it's hugely cheaper than proprietary models even with extra verbosity, and in his tests it drew a solid pelican-on-a-bike SVG, actually delivered a 5,000-word essay where most models quit early, and one-shot a Tuscan wellness-retreat homepage with images and animations — leading him to question paying for Sonnet or Gemini Flash tiers, while reminding viewers to check whether providers retain prompts for training. Run your own three-task probe on a new model — one visual/SVG task, one long-output task with a word target, one front-end build — and note where it quits early or exceeds expectations.

01

Brief

Start with this video's job: Sam Witteveen examines GLM 5.2 after Z.AI released both full and FP8 open weights, walking through its Artificial Analysis results (beaten only by GPT 5.5, Opus 4.8, and the withdrawn Fable 5), its long chain-of-thought token habits, $1.40/$4.40 per-million pricing, and hands-on tests from the pelican SVG to a 5,000-word essay and a design-arena-worthy homepage. Treat "Brief" as the outcome you are trying to make visible, not a topic label. Anchor it to 0:15, where the video says: “around with it this afternoon and the thing that just tipped me over the edge is artificial analysis have just published their stats on it as well. So one of the reasons I've been reluctant to make videos...”

02

Source material

Use "Source material" to locate the part of the creative automation mechanism the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 6:35, where the video says: “intelligence with short amounts of tokens in here. All right, the last one that I thought was really kind of interesting, too, was the design arena. So, they've got actually put this model at the front above sort...”

03

Generation

Turn "Generation" into the reusable artifact for this lesson: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints. This is where watching becomes something you can inspect and reuse.

04

Selection

Use "Selection" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.

05

Edit

Use "Edit" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.

06

Taste review

Use "Taste review" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.

07

Reusable recipe

Connect "Reusable recipe" to GLM 5.2 - The Top NEW Open Weights Model by naming the claim, the evidence, and the artifact it should produce.

Example

Source-backed artifact packet

Convert the video into a scoped artifact request that includes the transcript claim, mechanism, acceptance criteria, and proof. The output should be a creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints..

Example

Creative automation proof brief

Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the creative automation pattern.

Example

Teach-back module

Transform the lesson into a definition, a Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe diagram, one misconception, one practice exercise, and a check-for-understanding question.

Do not learn it wrong
  • Treating the title as the lesson without checking what the transcript actually says.
  • mistaking novelty for quality
  • no source/brief discipline
  • shipping generated media without taste review
  • Letting the lesson drift into generic content advice.
  • Letting the lesson drift into tool hype.
  • Letting the lesson drift into creative output without selection criteria.

Transcript-derived moments

Use timestamps to study the actual video.

Quality check

Do not count this as learned until these are true.

01

State the transcript-backed claim in your own words: Sam Witteveen examines GLM 5.2 after Z.AI released both full and FP8 open weights, walking through its Artificial Analysis results (beaten only by GPT 5.5, Opus 4.8, and the withdrawn Fable 5), its long chain-of-thought token habits, $1.40/$4.40 per-million pricing, and hands-on tests from the pelican SVG to a 5,000-word essay and a design-arena-worthy homepage.

02

Explain the practical stakes without hype: New playlist item from Sam Witteveen; queued for transcript-backed review, topic mapping, and a practical learning artifact.

03

Map the idea onto the Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe sequence and name the weakest link.

04

Produce the artifact and include the evidence that proves it: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.

Put it into practice

Give this grounded prompt to Codex or Claude after watching.

You are helping me turn one specific YouTube video into real, durable learning.

Source video:
- Title: GLM 5.2 - The Top NEW Open Weights Model
- URL: https://www.youtube.com/watch?v=10C8VMN3hjU
- Topic: Creative Automation
- My current learning frame: Test GLM 5.2 through OpenRouter on one task you currently pay a proprietary model for, measure tokens in and out at the $1.40/$4.40 rate versus your current cost, and check the provider's data-retention policy before deciding whether to switch.
- Why this matters: New playlist item from Sam Witteveen; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Transcript anchors from this exact video:
- 0:15 / Evidence 1: "around with it this afternoon and the thing that just tipped me over the edge is artificial analysis have just published their stats on it as well. So one of the reasons I've been reluctant to make videos..."
- 2:22 / Evidence 2: "more competitive with the Anthropic models and the OpenAI models, but this is a substantial bump from where they were for agentic coding than with the 5.1 model. So, the 5.1 model I thought was a a good..."
- 4:02 / Evidence 3: "just benchmark Fable without any fallback to 4.8, it actually does pretty badly because it has so many rejections for things that are often just total sort of nonsense rejections. And so if in cases like that, the..."
- 6:35 / Evidence 4: "intelligence with short amounts of tokens in here. All right, the last one that I thought was really kind of interesting, too, was the design arena. So, they've got actually put this model at the front above sort..."
- 8:12 / Evidence 5: "falling back and paying for the tokens that we actually use. Okay, so this is the tool that I use for basically testing a lot of models. We can do both local and served models in here. So,..."
- 10:08 / Evidence 6: "there. Okay, getting to the end, you can see that the word count for tokens did actually get above 5,000 tokens there. So, it certainly done a lot better than a lot of the other models for that..."
- 12:14 / Evidence 7: "check out. This could be an alternative to a lot of the proprietary models that you're using. I really don't see the point in using some of the models, perhaps like a Sonnet, perhaps like a Gemini Flash,..."

Video-aware target:
- Prompt lane: Creative automation
- Mechanism to extract: Extract the creative production loop, especially where the human keeps taste, selection, and final judgment.
- Artifact to produce: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.
- Artifact must include: brief; source inputs; generation recipe; selection criteria; edit/review checkpoint

Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, transcript support, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable mechanism from the video: Extract the creative production loop, especially where the human keeps taste, selection, and final judgment. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints.
5. Include:
   - a plain-English definition of the core idea
   - a diagram or structured model using this sequence: Brief -> Source material -> Generation -> Selection -> Edit -> Taste review -> Reusable recipe
   - answers to these source questions: What asset is being produced? | What inputs and tools drive it? | Where does human taste intervene?
   - 3 concrete examples that apply the video idea to real agentic work, such as Claude-generated video campaign; image-to-site workflow; voice or video editing loop
   - 2 failure modes the video helps prevent, chosen from the transcript evidence and these likely risks: mistaking novelty for quality; no source/brief discipline; shipping generated media without taste review
   - a checklist for the next real workflow, focused on: brief, inputs, generation, selection, critique
   - one practical exercise with a clear done signal: Build one reusable creative recipe and define what would make the result rejectable.
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.

Quality bar:
- Make this specific to "GLM 5.2 - The Top NEW Open Weights Model", not a generic Creative Automation essay.
- Anchor each creative step to transcript evidence about inputs, model/tool choices, iteration, editing, or critique.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- Avoid these generic drifts: generic content advice; tool hype; creative output without selection criteria.
- If evidence is weak or missing, stop and say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.

Misconceptions

What to stop believing.

Creative AI removes the need for taste.

It increases the need for taste because output volume explodes.

The best prompt is enough.

References, critique, iteration, and post-production matter just as much.

Practice studio

Learning only counts when you make something.

01

Transcript evidence map

Separate what the video actually says from what you already believe about the topic.

3 source-backed takeaways with timestamps, confidence, and a transfer note.
02

One useful artifact

Apply the video to a real workflow and produce a creative production board with source inputs, prompt recipe, selection criteria, edit pass, and taste-review checkpoints..

A reusable artifact with a done signal and one verification step.
03

Creative automation teach-back card

Explain the creative automation mechanism to someone who has not watched the video yet.

A 90-second explanation, one diagram, one example, and one misconception to avoid.

Recall check

Answer first, then reveal — without rewatching.

What did Z.AI release for GLM 5.2, and why does Witteveen consider withholding base models fair?

What caveat does the video attach to GLM 5.2's strong Artificial Analysis score?

What is GLM 5.2's API pricing, and which hands-on results impressed Witteveen most?

Source shelf

Use the video as a doorway, then verify with primary sources.

ReadingComfyUIwww.comfy.org/ReadingAffinityaffinity.serif.com/