“I used Codex to make Seedance 2.5 arrive at the Disaster Girl frame”
ChatGPT is usually treated as a prompt box.
Ask for a cinematic shot, add a few camera words, copy the result into a video model, and hope the ending survives.
My problem was that the ending usually did not survive.Most of my image-to-video failures happen in the last second. The motion can look coherent for 19 seconds, then Seedance suddenly snaps into the reference image, as if it remembered the endpoint too late. Faces morph, perspective jumps, or a white flash hides the transition.
So I changed the role of ChatGPT.Instead of asking it to write a more impressive prompt, I used ChatGPT/Codex as a shot planner. The final Disaster Girl frame came first. The whole camera route had to be designed backward from that destination.
The useful shift was:
stop asking “what should happen in this video?”
start asking “what has to be true when the video ends?”
1.Give ChatGPT the terminal state
I started by listing everything that could not change in the final frame:
the girl’s position, gaze and expression, the foreground/background order, the final camera height, and the exact viewing axis.Those became constraints rather than visual suggestions.
2.Separate camera actions from visual adjectives
Words like “dynamic,” “smooth,” and “cinematic” sound useful, but they do not define a route.I asked Codex to turn the final-frame requirements into a shot plan with explicit camera movement, foreground events, exposure changes and stopping behavior.The goal was not a longer prompt. The goal was to remove ambiguity.
3.Backsolve a physical route
The resulting route looked like this:
low interior start → follow a red cue across the floor → accelerate toward a blown-out exit → cross into the yard → pass the drill crew, hose lines and fire truck as foreground occlusions → move alongside the girl → arc onto the final viewing axis
Each object had a reason to exist.The red cue establishes direction. The bright exit motivates the exposure transition. The truck and hose lines create translation and parallax. The final arc brings the camera onto the reference perspective.
4.Use ChatGPT to find the failure point
When the first version drifted, I did not ask for “a better prompt.”I asked which constraint was underspecified.
That usually exposed a specific problem:
the camera was rotating instead of translating, the foreground objects had no timing, or the final frame was described as an image instead of a state the camera had to reach.
That made iteration much more mechanical.
5.Write the stop as an action
The ending became its own motion phase:
fast flight → controlled glide → small positional correction → shared camera/subject deceleration → complete stop
The camera and facial movement settle together on the final frame.The stop is not an edit. It is the last action in the shot.
For reproducibility: the opening frame was generated in Seedream 5.0 Pro for $0.045. The 20-second Seedance 2.5 run cost $2.68 at 1080p-ESR / 60 fps through the Atlas Cloud API inside OpenMontage.
I also referenced this repo while turning the reverse-engineered route into a model-ready shot plan:
The interesting part for me was not that ChatGPT wrote a good prompt. It was that it helped turn a vague visual idea into a set of constraints, causes and transitions that another model could execute.
Has anyone else used ChatGPT as a shot planner instead of a prompt rewriter? I’m curious whether the same terminal-state approach would work for UI animation, robotics simulations, architecture, or any task where the final state matters more than the first frame.
r/ChatGPT3.7K