JuicyBite, part 4 of 7

Giving a brand a story without a studio

For a restaurant nobody walks into, the story is the storefront.

2026Operations and business analyst intern

Context

A delivery restaurant has nowhere to put its brand. There’s no sign, no dining room, no staff to remember. The whole impression is a name and a photo in an app. For Hidden Sichuan, I wanted more than that: a story world people could follow before they ever ordered.

The problem

The usual way to get an animated series is to hire a studio, and for a business that hadn’t opened yet, that wasn’t realistic. I decided to make it myself with AI image and video tools.

The tools turned out to be the easy part. The hard part was getting them to produce the same characters, the same kitchen and the same food, consistently, across dozens of shots.

What I did

Building the world first

Before generating anything, I wrote the world: who the characters are, where the story happens, and who they’re up against. That included a rival brand, Double Dragon, written as Hidden Sichuan’s opposite.

Managing only what repeats

My first list of visual assets had 104 items. I cut it to 87, then to 51. The rule: manage only what appears again and again. Anything that shows up once gets described directly when it’s needed.

A character guide someone else could use

I made a 36-sheet character guide that an outside artist could work from, with every character shown from several angles. I designed each character’s body as a record of their working life rather than a default “attractive” shape, and gave each group in the story its own color system.

An episode built on a beat

The first episode is 2 minutes and 16 seconds long: 68 shots, every one exactly two seconds, cut to a 120 beats-per-minute track. At its center is a 60-second recipe sequence in which each of the six cooking stages adds a new layer to the music, from the kick drum to the melody, with the first kick timed to the moment the flame ignites.

Separately, I generated a full cooking sequence of 86 shots, first as still images and then as video clips.

Doing the research

Content only feels real if it’s accurate. I wrote 39 pages of research notes on mala: the six layers of flavor in the sauce, where each ingredient comes from, the Chinese national standards for key ingredients, the science of fermentation, and food safety. It ends with a list of 20 claims that sound true but shouldn’t be stated as fact.

Making the process cheaper

Each failed video costs far more to redo than a still image. So I locked the key frames as images first, and only then generated video from them. I also built a simple tool for locking the visual style: about 75 references across nine decisions, with a checker that flagged problem words in prompts.

What I learned

AI tools reward precision more than imagination. When I described a pile of food as “crown-shaped,” the model drew an actual crown in the pan. It happened with wings, ribbons and domes too. Once I banned comparisons and described shape and physics directly, wider at the base and narrowing toward the top, the results became consistent. It wasn’t a creative problem. It was a specification problem.

I also learned that cost structure decides workflow. Knowing what was cheap to redo and what was expensive changed the order I did everything in.

And a day after producing 86 clips, I realized that a few months earlier this would have started with a request for a quote from a studio. Knowing how the tools work is a real gap between people now.