Creative Automation / Foundation

I Built a Vox-Style Video Using HyperFrame and Claude Code

Andy Lo builds a complete Vox-style editorial animation about China's EV and battery dominance using HyperFrames β€” HeyGen's open-source framework that lets AI agents create videos as HTML/CSS/JavaScript β€” driven by Claude Code through a three-phase prompt workflow (foundation and design system, scene building with asset review, then narration/sourcing/render), then levels it up from seven scenes to 19 subscenes with aligned narration.

Andy Lo21 minTranscript found

Quick learning frame

Read this before watching.

A design-system lesson is about making visual taste reusable through tokens, components, examples, constraints, and review loops.

New playlist item from Andy Lo; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Skill you build: The ability to run code-based video production with an AI agent β€” structuring a storyboard brief, CLAUDE.md rules, fonts, and an asset library so Claude Code builds a cohesive motion-design system, and iterating via review passes and handoff files rather than frame-by-frame editing.

Watch for the shift from claim to mechanism. The learning value is the point where the transcript reveals a repeatable action, tool boundary, context move, review habit, or artifact.

Concept diagram

Where this video fits.

01Reference
02Tokens
03Components
04Usage rules
05Agent prompt context
06Implementation
07Visual QA

Deep lesson

Turn this video into working knowledge.

3,647 cleaned transcript words reviewed across 1,110 timed caption segments.

Thesis

I Built a Vox-Style Video Using HyperFrame and Claude Code teaches a practical design system move: Andy Lo builds a complete Vox-style editorial animation about China's EV and battery dominance using HyperFrames β€” HeyGen's open-source framework that lets AI agents create videos as HTML/CSS/JavaScript β€” driven by Claude Code through a three-phase prompt workflow (foundation and design system, scene building with asset review, then narration/sourcing/render), then levels it up from seven scenes to 19 subscenes with aligned narration.

The goal is not to remember the video. The goal is to extract the operating principle, tie it to timestamped evidence, test how far the claim transfers, and make something reusable.

1:08

Video as a web project

β€œAI video avatars and realistic lip sync. But hybrid frames is pretty different. Hyper frames is an open-source video framework from hijen that lets AI agents create videos using HTML, CSS and JavaScript. So in simple terms, it...”

HyperFrames (from HeyGen, better known for avatars and lip sync) turns video creation into code: Claude Code builds the video like a web project in HTML, CSS, and JavaScript and HyperFrames renders it β€” solving the problem that prior AI motion graphics (like Claude with Remotion) felt like animated slides, by giving control over pacing, layout, timing, typography, and transitions; the whole workflow is three stages: foundation, story (copy, scenes, GSAP animation, transitions), and rendering with narration and citations. Write out the three-stage pipeline (foundation, story, render) and under each stage list the artifacts it produces, so you can map any video idea onto the structure before prompting.

5:52

Foundation before scenes

β€œtrying to create. And in our case, that's a fed style video about electric vehicles and artificial intelligence. And once Claude understands the goal, it starts building the design system. And that includes things like the typography, color...”

Prompt one makes Claude read the storyboard-and-direction brief first, then build the design system β€” typography, color palette, and a motion language taught to mimic Vox's intentional, measured, slightly choppy editorial style β€” plus placeholders for all seven scenes so nothing looks inconsistent later; the ~20-minute result is deliberately bare-bones with correct hierarchy and motion, and prompt two then turns collected assets into full scenes with the rule that every visual must explain something and ends by flagging weak or missing assets for review. Draft a one-page storyboard-and-direction brief for a topic you know: story, visual style, scene list, palette, typography, and the one message viewers should walk away with.

15:15

Inputs decide quality

β€œinto the timeline. And now another thing you may have probably noticed is that we have had a open code session running in a separate tab this entire time. And we actually used open code together with all...”

The leveled-up version keeps the exact same workflow but upgrades the inputs β€” seven chapters split into 19 subscenes for denser pacing, over 70% of the asset library used, and a narration script pre-aligned to the timeline (with ElevenLabs voice wired up via OpenCode) β€” producing a video that feels publishable rather than prototype; the lesson is that tools are only part of the equation: better storyboards, assets, and narration, i.e. creative direction, drive the final quality, and HyperFrames amplifies motion designers rather than replacing them. Take one finished AI-generated piece of yours and upgrade only the inputs β€” a more granular outline and more/better source assets β€” without changing tools, then compare the two outputs.

01

Reference

Start with this video's job: Andy Lo builds a complete Vox-style editorial animation about China's EV and battery dominance using HyperFrames β€” HeyGen's open-source framework that lets AI agents create videos as HTML/CSS/JavaScript β€” driven by Claude Code through a three-phase prompt workflow (foundation and design system, scene building with asset review, then narration/sourcing/render), then levels it up from seven scenes to 19 subscenes with aligned narration. Treat "Reference" as the outcome you are trying to make visible, not a topic label. Anchor it to 1:08, where the video says: β€œAI video avatars and realistic lip sync. But hybrid frames is pretty different. Hyper frames is an open-source video framework from hijen that lets AI agents create videos using HTML, CSS and JavaScript. So in simple terms, it...”

02

Tokens

Use "Tokens" to locate the part of the design system mechanism the video is demonstrating. Ask what changes in your real setup if this claim is true. Anchor it to 5:52, where the video says: β€œtrying to create. And in our case, that's a fed style video about electric vehicles and artificial intelligence. And once Claude understands the goal, it starts building the design system. And that includes things like the typography, color...”

03

Components

Turn "Components" into the reusable artifact for this lesson: A design-system adoption brief with source references, tokens/components, agent handoff rules, and visual QA checks. This is where watching becomes something you can inspect and reuse.

04

Usage rules

Use "Usage rules" as the application surface. Decide whether the idea touches a browser flow, a local file, a model choice, a source document, a UI, or a review step.

05

Agent prompt context

Use "Agent prompt context" to prove the lesson. The evidence should connect back to the video title, transcript anchors, and a concrete output, not a generic best-practice claim.

06

Implementation

Use "Implementation" to carry the idea forward: save the prompt, checklist, diagram, or operating rule that would make the next agent run better.

07

Visual QA

Connect "Visual QA" to I Built a Vox-Style Video Using HyperFrame and Claude Code by naming the claim, the evidence, and the artifact it should produce.

Example

Source-backed artifact packet

Convert the video into a scoped artifact request that includes the transcript claim, mechanism, acceptance criteria, and proof. The output should be a design-system adoption brief with source references, tokens/components, agent handoff rules, and visual qa checks..

Example

Design system proof brief

Separate what the speaker claims, what the demo actually proves, and what still needs outside verification before you adopt the design system pattern.

Example

Teach-back module

Transform the lesson into a definition, a Reference -> Tokens -> Components -> Usage rules -> Agent prompt context -> Implementation -> Visual QA diagram, one misconception, one practice exercise, and a check-for-understanding question.

Do not learn it wrong
  • Treating the title as the lesson without checking what the transcript actually says.
  • copying visuals without rules
  • generic generated UI
  • no visual QA screenshot pass
  • Letting the lesson drift into generic design inspiration.
  • Letting the lesson drift into component lists without usage rules.
  • Letting the lesson drift into no screenshot review.

Transcript-derived moments

Use timestamps to study the actual video.

Quality check

Do not count this as learned until these are true.

01

State the transcript-backed claim in your own words: Andy Lo builds a complete Vox-style editorial animation about China's EV and battery dominance using HyperFrames β€” HeyGen's open-source framework that lets AI agents create videos as HTML/CSS/JavaScript β€” driven by Claude Code through a three-phase prompt workflow (foundation and design system, scene building with asset review, then narration/sourcing/render), then levels it up from seven scenes to 19 subscenes with aligned narration.

02

Explain the practical stakes without hype: New playlist item from Andy Lo; queued for transcript-backed review, topic mapping, and a practical learning artifact.

03

Map the idea onto the Reference -> Tokens -> Components -> Usage rules -> Agent prompt context -> Implementation -> Visual QA sequence and name the weakest link.

04

Produce the artifact and include the evidence that proves it: A design-system adoption brief with source references, tokens/components, agent handoff rules, and visual QA checks.

Put it into practice

Give this grounded prompt to Codex or Claude after watching.

You are helping me turn one specific YouTube video into real, durable learning.

Source video:
- Title: I Built a Vox-Style Video Using HyperFrame and Claude Code
- URL: https://www.youtube.com/watch?v=XVsGK99E9FA
- Topic: Creative Automation
- My current learning frame: Pick a topic you understand well, write a storyboard-and-direction brief plus a CLAUDE.md of project rules, gather a generous asset folder, then run the three-prompt HyperFrames workflow with Claude Code β€” foundation, scenes with a weak-asset review pass, optional narration β€” and iterate until the motion language feels editorial rather than slide-like.
- Why this matters: New playlist item from Andy Lo; queued for transcript-backed review, topic mapping, and a practical learning artifact.

Transcript anchors from this exact video:
- 1:08 / Evidence 1: "AI video avatars and realistic lip sync. But hybrid frames is pretty different. Hyper frames is an open-source video framework from hijen that lets AI agents create videos using HTML, CSS and JavaScript. So in simple terms, it..."
- 2:49 / Evidence 2: "inspired design language, the animation style, the scene structure, the color palette, typography choices, and the key message we want the audience to walk away with. So ultimately what this document does is give Claude a clear understanding..."
- 5:52 / Evidence 3: "trying to create. And in our case, that's a fed style video about electric vehicles and artificial intelligence. And once Claude understands the goal, it starts building the design system. And that includes things like the typography, color..."
- 8:38 / Evidence 4: "asking it to transform those raw assets into seven fully built scenes based on our storyboard. Okay. So now notice that we're not just telling Claude what assets to use. We are also reminding it how to use..."
- 11:32 / Evidence 5: "And as the quotes builds, it will essentially autoco compact its context. And this allows you to continue working with a fresh context window without starting over from scratch. And for most projects that's perfectly fine. But for..."
- 15:15 / Evidence 6: "into the timeline. And now another thing you may have probably noticed is that we have had a open code session running in a separate tab this entire time. And we actually used open code together with all..."
- 20:37 / Evidence 7: "frames. So instead of editing every frame manually, you just describe what you want and generate the system that creates it. That means faster iteration, better consistency, and workflows that are dramatically easier to scale. Try this yourself."

Video-aware target:
- Prompt lane: Design system
- Mechanism to extract: Extract how the video turns visual references or component systems into usable constraints for agents.
- Artifact to produce: A design-system adoption brief with source references, tokens/components, agent handoff rules, and visual QA checks.
- Artifact must include: references; tokens/components; handoff artifact; implementation rule; visual QA

Your task:
1. Use the transcript anchors above as the primary source packet. If you add outside context, label it clearly as outside context and keep it secondary.
2. Create a source-check table with columns: timestamp, claim, transcript support, what the demo proves, confidence, and what still needs verification.
3. Extract the actual teachable mechanism from the video: Extract how the video turns visual references or component systems into usable constraints for agents. Do not invent claims that are not supported by the title, lesson frame, or transcript anchors.
4. Build a reusable learning artifact: A design-system adoption brief with source references, tokens/components, agent handoff rules, and visual QA checks.
5. Include:
   - a plain-English definition of the core idea
   - a diagram or structured model using this sequence: Reference -> Tokens -> Components -> Usage rules -> Agent prompt context -> Implementation -> Visual QA
   - answers to these source questions: What design source is reused? | How is it translated into agent context? | What review catches generic output?
   - 3 concrete examples that apply the video idea to real agentic work, such as Figma-to-shadcn workflow; design.md brief; UI reference library remix
   - 2 failure modes the video helps prevent, chosen from the transcript evidence and these likely risks: copying visuals without rules; generic generated UI; no visual QA screenshot pass
   - a checklist for the next real workflow, focused on: references, tokens, components, handoff, QA
   - one practical exercise with a clear done signal: Turn one screen reference into five constraints a coding agent must follow.
6. Add a "learning transfer" section: what changes in my workflow tomorrow if I actually learned this?
7. Add a "source check" section that cites which transcript anchor supports each major takeaway.

Quality bar:
- Make this specific to "I Built a Vox-Style Video Using HyperFrame and Claude Code", not a generic Creative Automation essay.
- Cite transcript anchors for every claim about design context, UI generation, preview, critique, or handoff.
- Prefer operational examples, failure modes, and reusable artifacts over broad definitions.
- Call out uncertainty instead of smoothing over weak evidence.
- Avoid these generic drifts: generic design inspiration; component lists without usage rules; no screenshot review.
- If evidence is weak or missing, stop and say what transcript segment or timestamp needs review instead of guessing.
- Finish with a concise artifact I could paste into my learning app.

Misconceptions

What to stop believing.

Creative AI removes the need for taste.

It increases the need for taste because output volume explodes.

The best prompt is enough.

References, critique, iteration, and post-production matter just as much.

Practice studio

Learning only counts when you make something.

01

Transcript evidence map

Separate what the video actually says from what you already believe about the topic.

3 source-backed takeaways with timestamps, confidence, and a transfer note.
02

One useful artifact

Apply the video to a real workflow and produce a design-system adoption brief with source references, tokens/components, agent handoff rules, and visual qa checks..

A reusable artifact with a done signal and one verification step.
03

Design system teach-back card

Explain the design system mechanism to someone who has not watched the video yet.

A 90-second explanation, one diagram, one example, and one misconception to avoid.

Recall check

Answer first, then reveal β€” without rewatching.

What is HyperFrames and how does it differ from HeyGen's better-known products?

What does prompt one accomplish before any scenes are built, and why?

What changed between the first build and the 'leveled up' version, and what lesson does Andy draw?

Source shelf

Use the video as a doorway, then verify with primary sources.

ReadingComfyUIwww.comfy.org/ReadingAffinityaffinity.serif.com/