I Couldn't Learn After Effects, So I Built My Own. It Taught Me Product.
The problem
Here's a confession: I love motion graphics, and I could never learn
After Effects.
I wanted to make catchy short videos. A stat that slams onto the screen.
A lyric video. A quick promo for my own app. Every time, I'd open After
Effects, stare at a wall of panels, layers and graph editors, watch a
tutorial or three, and close it with nothing made.
I had the ideas. I just couldn't turn them into motion. So instead of
learning After Effects, I built my own motion graphics studio, where an
AI agent does the keyframes and I do the taste. What I didn't expect was
that building it would teach me more about product than about animation.
Not another "type a prompt, get a video" app
The obvious move was a text-to-video model. I didn't want one. Generated
video is a slot machine: you pull the lever, get something close, and
then you can't say "move that word up 20 pixels." You just regenerate
and hope.
So instead, the agent operates a real animation engine. Every video is
one scene file with layers, keyframes and a camera. The agent edits that
file using the exact same actions I use in the UI. That means everything
it makes is editable, frame accurate and consistent from one clip to the
next. If one word is off, I fix one word, and nothing else moves.
The agent isn't the renderer. It's the crew.
The deal: I judge, it executes
I split the work by what each side is actually good at. I'm good at
deciding what a video is about, pointing at what matters, picking
between options and reacting to a preview. I'm bad at (and bored by)
easing curves, layout math and timing forty layers by hand. The agent is
the opposite.
So I talk to it in a chat panel, it builds the scene live, and every
turn it takes is one undo step. There's also a Direct mode: I drag the
camera around roughly, like waving my phone, and the agent cleans it up
into a smooth, precise move.
One rule I'm proud of: user edits are sacred. If I touch something,
the agent can't overwrite it unless I explicitly say so. That's enforced
in the code, not just asked nicely in a prompt.
Teaching taste with rules
"Make it catchy" isn't an instruction anyone can follow, human or AI. So
I wrote down what catchy actually means: hook the viewer in the first
second, overshoot and settle, stagger things, keep something always
moving, mix fast and slow, and leave text up long enough to read.
Then I gave the agent a way to check itself. After it builds something,
it renders a contact sheet of its own frames and asks four questions:
- Is the first frame dead?
- Is anything sitting still too long?
- Can every line be read in time?
- Is too much happening at once?
Good motion doesn't come from hoping the model has taste. It comes
from rules plus a mirror.
Then it got a little out of hand
Once one agent could build a video, I wondered what would happen if I
let a few of them loose. Now I have a newsroom: six agents, two each for
YouTube Shorts, Instagram Reels and TikTok.
Each one picks its own niche, researches facts on the web, pitches an
idea with sources, builds the video, checks its own frames and drops the
finished clip into a Slack-style channel for me to review. They've come
up with titles like "$7.25 hasn't moved in 17 years" and "A Day on
Mercury Lasts Two of Its Years."
So far they've made 25 videos at about $1.70 of model usage each, all
running on my Claude Max plan. A few things stay human on purpose: every
fact needs a source, the agents only use material we own, and I'm the
one who hits post.
The product calls behind it
The code was the easy part to delegate. The hard part was deciding what
to build, what to skip and what each choice would cost. Here are the
calls that shaped the studio most.
Bring it all into one app
My first plan was modest. I'd learn just the slice of After Effects I
needed, and let the studio make graphics while DaVinci Resolve handled
the rest. But every hop between apps cost time, so one by one the pieces
moved in: MP4 export, music, footage, data charts. Today the studio
assembles almost the whole video, audio included. Keeping everything in
one place simply made more sense than stitching tools together.
That's when it clicked. After Effects probably went through the exact
same thing. Every feature makes sense on its own, and together they add
up to the wall of panels that scared me off in the first place.
The trade-off: every feature I add makes my studio a little more
like the thing I couldn't learn. Keeping it simple is a decision I have
to keep making, not one I made once.
Let a budget shape the architecture
I wanted it to run on my flat-rate Claude Max plan, not a pay-per-token
API key. It turned out a subscription login only works through Claude
Code itself. So instead of fighting that, I designed around it: the
studio launches Claude Code behind the scenes and hands it the studio's
tools, while the chat stays inside the app.
The trade-off: a slightly unusual setup. In return, every video
costs me nothing extra beyond the plan I already pay for.
Remove a limit that protected nothing
Early on, the agent had a cap on how many steps it could take per
request, and long builds kept dying halfway. I asked myself why the cap
existed at all. On a flat-rate plan it saved no money, and it threw away
finished work. So I removed it by default and added a Stop button
instead.
The trade-off: a runaway turn is possible, so I'm the one who stops
it. In return, the agent finishes what it starts.
One agent per video, not a team per video
My first plan was a lead agent splitting one video into sections for
several workers. Then I looked at what I actually make: short videos.
Splitting a 30-second clip adds seams, a shared camera and messy undo
for almost no speed gain. Running several agents on separate videos
gives the same output with none of that.
The trade-off: a single long video doesn't get faster. In return,
the system stays simple.
Do the math before buying the feature
I wanted music on every video. My music plan gives about 33 minutes of
generation a month, and six agents making two videos a day would need
around 180. Fresh music per video was never going to work. So I
generated a library of 16 tracks once, all built with the same
structure, and each video borrows the stretch that ends on the track's
own outro. Then I cut scope again: only YouTube gets library music,
because on Instagram and TikTok I add trending sounds myself for reach.
The trade-off: less variety than fresh tracks. In return, music fits
the budget and actually helps where it's used.
Keep the risky parts human
The agents can research, build and render on their own, but they can't
post. Auto-posting would mean developer apps and review processes on
every platform, and one bad autonomous post is a brand risk I don't
want. Every fact also needs a source, and agents only use material we
own.
The trade-off: I'm still in the loop for every post. That's the
point.
Read the license before the README
I turned down a strong video-tracking model because its license didn't
allow commercial use, and used a different approach instead. I also
replaced several libraries with small in-house code to keep dependencies
down, and noted one license (in beat detection) that would need a second
look if I ever sell the studio. Licenses are product decisions, not
paperwork.
What surprised me (and what broke)
- The agents all had the same idea. All six picked "by the numbers"
niches, and three of them landed on money and prices. I'd only told
partners on the same platform to stay different from each other.
Lesson: if you want variety, it has to be a rule for everyone.
- The toolbox shapes the content. The chart tools were the
strongest, so the agents made chart videos. What you build into the
tools quietly decides what gets made.
- I'm the bottleneck now. All 25 videos are sitting in my review
queue. The work moved from making to judging, which is exactly what I
wanted, and also a little humbling.
- The agents are decent coworkers. One left a note saying the critic
seemed to mishandle multi-line text. It was right, and I fixed it.
Real use also broke things that tests never did. When my laptop went to
sleep, the agent's background browser died silently and previews froze.
And some of the generated music ended with up to three seconds of
silence, which would have ended videos on dead air. Both are fixed now.
The part I can't automate
I've started posting these videos on YouTube, Instagram and TikTok. Part
of it is a real experiment: will anyone be able to tell that AI did most
of the work? Part of it is just wanting the channels to grow. I don't
have numbers worth sharing yet, so check back in a few months.
But I've already learned the uncomfortable lesson. Building the studio
was the easy part. Getting people to actually see what it makes is the
hard part, and I'm not sure anyone ever fully solves it. Coding can be
delegated. Distribution, so far, can't.
What I learned
I built this in about two weeks with Claude Code writing most of the
code. My job was the design, the rules, the trade-off calls and actually
using the thing until it broke. Some of the best features came from
frustration, like autosave, which showed up right after I asked, "so the
previous work never got saved?"
The biggest surprise was After Effects itself. I started out thinking it
was just overcomplicated. After watching my own tool grow one sensible
feature at a time, I get it now. Every panel that intimidated me was
probably someone's very reasonable product decision.
The biggest lesson: AI didn't remove the work, it moved it. I still
don't know After Effects, and now I don't need to. My time goes into
deciding what's worth making.
Honestly, that's the part I wanted to be doing all along.
← All posts