AI Video12 min

How to Keep a Character Consistent From Reference Frame to Video

Build a character reference in Nano Banana Pro, choose the right first-frame, paired-frame, or reference-to-video route, then review identity at the first, hardest, and last frames.

Yingtu AI Editorial
Yingtu AI Editorial
YingTu Editorial
Jun 15, 2026
Updated Jul 28, 2026
12 min
A character continuity workflow moving from an approved Nano Banana Pro reference frame through video route selection to first, middle, and last-frame review.
yingtu.ai

Contents

No headings detected

The most reliable way to keep a character consistent in image-to-video is to stop treating “same character” as a prompt phrase. Make it a production contract: decide which traits are locked, prepare reference evidence for the hardest shot, animate one short diagnostic clip, and inspect the character at the first frame, the most demanding motion frame, and the last frame.

Nano Banana Pro belongs at the still-image and reference-frame layer in this workflow. Use it to create or repair an approved anchor, profile, full-body view, outfit detail, or ending composition. A separate video route creates the motion. No reference method guarantees zero drift, so the output still needs frame-level review.

Start with one rule: a clip does not pass just because frame one looks correct. Hair can change under motion, body proportions can stretch during a turn, a signature prop can disappear behind an occlusion, or the visual style can shift before the final frame.

Choose the reference method before you rewrite the prompt

The input route decides what evidence the video model receives. Use the lightest route that contains the information your hardest required shot needs.

Input methodUse it whenWhat it actually controlsDo not assume
Single first frameOne short shot begins close to the approved composition, with modest camera or character movementThe visible subject, outfit, lighting, style, and composition at frame zeroA front portrait will remain sufficient through a profile turn, full-body action, or heavy occlusion
Paired first and last framesThe approved ending composition matters: a reveal, entrance, pose change, product handoff, or match cutThe two endpoints supplied to a route that supports both imagesCorrect endpoints guarantee correct identity or motion between them
Reference-to-videoSubject identity, outfit, prop, or style evidence must remain available beyond frame zeroRoute-specific content or subject references that can guide the generated clipMore references always help, or every provider maps them the same way
Structured or custom routeThe production needs unusual angles, repeated multi-shot identity, controlled poses, or a critical likeness that lighter conditioning repeatedly missesStronger identity or structural conditioning, depending on the selected systemTraining or extra controls remove the need for rights review, output QA, or manual cleanup

Google's current Veo 3.1 documentation distinguishes these modes: an initial image can become the first frame, a lastFrame can define an interpolation endpoint, and up to three reference images can guide content on supported routes. That is evidence that the input modes exist—not proof that they preserve every identity detail. Check the current Gemini API video documentation for the selected route before production.

If you cannot name the hardest shot yet, do not select a route. A close-up talking head, a full-body run, a rear three-quarter turn, and a two-character handoff expose different missing evidence.

Build evidence for the hardest shot in Nano Banana Pro

Do not begin with a large gallery of nearly identical portraits. Begin with one approved anchor, then add only the views that resolve real uncertainty.

For a fictional mascot crossing a room and turning toward camera, a useful reference pack might contain:

  • one three-quarter anchor with the approved face, hair silhouette, body proportions, outfit, prop, and rendering style;
  • one profile view because the turn exposes the nose, jaw, ear, and hairline;
  • one full-body view because the walk tests height, limb proportions, shoes, and costume length;
  • one clean prop detail if the prop is story-critical;
  • an approved final composition only if the selected route accepts a last frame and that ending matters.

Nano Banana Pro can help create or repair those still references, but approve them before they enter video generation. If two images disagree about jaw width, hair length, jacket seams, or prop scale, the video route receives a contradiction rather than extra evidence.

Create a lock/change matrix beside the references:

Lock for this shotAllowed to change deliberately
Face geometry and identifying featuresExpression and gaze
Hair silhouette, length, and colorHair motion caused by wind or movement
Body proportions and apparent agePose and action
Signature outfit, accessories, and propsCamera angle and framing
Rendering style, line quality, texture, and color logicLighting and setting, within the approved art direction

This matrix keeps a useful motion request from fighting the identity request. “Run toward camera while the coat moves” is compatible with a locked coat design; “change into a different outfit” is not.

Match reference framing to the shot

A clear portrait is good evidence for a close-up, but weak evidence for a full-body spin. The LTX character-consistency guide makes this framing relationship explicit: the reference should resemble the intended starting shot, and multi-shot work benefits from a reference library with varied angles and framings.

Use this practical test: hide every reference that does not show the body part, angle, outfit detail, or prop that the hardest moment will reveal. If the remaining evidence is thin, repair the reference pack before spending video generations.

Run one hard diagnostic clip

The first test is not a beauty shot. It is a controlled attempt to discover whether the selected route can preserve the required identity under the hardest meaningful condition.

  1. Pick one short shot that contains one identity challenge: a head turn, full-body step, brief occlusion, or change from medium to wide framing.
  2. Use only approved, rights-cleared references.
  3. Ask for one major motion or camera change, not five.
  4. Keep the locked traits explicit and the allowed changes narrow.
  5. Generate the shortest clip that reaches the demanding moment.
  6. Save the input version, route, model, date, prompt, and output together.
  7. Review the first frame, the demanding middle frame, and the last frame before deciding what to change.

A diagnostic prompt can be compact:

Medium shot. The mascot takes two steps forward and turns 45 degrees toward camera. Keep face geometry, short black hair silhouette, body proportions, teal jacket with two silver fasteners, red satchel, and flat cel-shaded line style unchanged. Natural cloth movement only. No new accessories, no outfit change, no camera cut.

The reference supplies appearance. The prompt specifies motion, camera, timing, locked traits, and forbidden drift. Repeating a long biography does not replace missing visual evidence.

Review five identity dimensions at three frames

Use the same five dimensions at the beginning, the hardest motion moment, and the end. Record a simple pass, repair, or critical fail for each cell, with a short observation.

Identity dimensionFirst frameDemanding motion frameLast frameWhat counts as drift
FaceGeometry, age, feature spacing, identifying mark, or expression changing beyond the allowed performance
HairSilhouette, length, texture, part, or color changing rather than moving naturally
BodyHeight impression, limb length, shoulder/hip proportions, or apparent age changing
Outfit and propsGarment construction, color, accessory placement, logo, or signature prop changing or disappearing
Visual languageLine weight, texture, materials, color logic, realism level, or character-to-background rendering style shifting

Add three shot-level checks below the identity grid:

  • Motion: Is the requested action readable without anatomy or prop failure?
  • Camera: Did the framing and movement follow the shot plan?
  • Ending composition: If an ending was specified, did the clip reach it without sacrificing a critical identity trait?

Set the acceptance rule before review. A reasonable fictional-mascot rule might be:

  • every critical identity dimension must pass at all three checkpoints;
  • one non-critical outfit texture variation may be repairable if it is not visible at delivery size;
  • a face, body-proportion, or signature-prop failure is never averaged away by four easier passes;
  • the hardest required shot—not the easiest clip—decides whether the route is viable.

This is a project threshold, not a universal benchmark. Ads may require exact product and mascot details; an impressionistic short may allow more texture variation; a client avatar may require a stricter consent and likeness boundary than either.

Repair the smallest true failure

Blind rerolling hides the cause. Use the visible failure to choose the smallest next move.

hljs text
Did a critical identity dimension fail?
├─ No
│  └─ Did motion, camera, timing, or ending composition fail?
│     ├─ Yes → revise only that motion variable and rerun the same reference setup
│     └─ No  → approve the clip and its final frame for the next handoff decision
└─ Yes
   ├─ Is the needed identity evidence missing or contradictory in the still pack?
   │  ├─ Yes → repair the Nano Banana Pro still or replace the conflicting reference
   │  └─ No
   ├─ Does the reference framing mismatch the hardest moment?
   │  ├─ Yes → supply the required profile, full-body, outfit, or prop view
   │  └─ No
   ├─ Is the shot asking for too much camera/action change at once?
   │  ├─ Yes → shorten the clip and remove all but one major change
   │  └─ No
   ├─ Would missing endpoint or persistent identity evidence solve the diagnosed gap?
   │  ├─ Yes → switch from first-frame to paired-frame or reference-to-video
   │  └─ No  → evaluate a stronger structured/custom route or budget manual cleanup
   └─ Same critical dimension fails twice in reduced-variable tests
      → stop spending on that setup; switch route, revise the shot, or stop the asset

That final stop is important. If the same jaw shape, full-body proportion, or signature prop fails twice after you have reduced motion and supplied matching evidence, a longer prompt is not a diagnosis. Record the failure and change the production decision.

Common failures and minimum fixes

Visible failureLikely evidence gapMinimum next test
Face changes only during a profile turnReference pack contains only frontal evidenceAdd one approved profile or three-quarter reference; keep the same short turn
Coat, logo, or prop mutates under fast actionToo much motion plus a small or ambiguous design detailUse a clearer still, enlarge the key detail, and reduce the action
Character matches at both endpoints but changes in betweenPaired frames constrain composition more than temporal identityReduce motion or test reference-to-video/stronger conditioning; do not call the clip passed
Correct character, wrong endingSingle first frame lacks endpoint evidenceUse a supported paired-frame route if the ending is essential
Every reroll fails a different wayConflicting references or too many simultaneous variablesReturn to one approved anchor and one hard shot
Clip is acceptable, but export/rights are unusableProvider or route contract failureSwitch route; prompting cannot repair ownership, watermark, or rights terms

Approve every cross-clip handoff

Using the previous clip's last frame as the next clip's first frame can improve visual continuity, but only if that frame has already passed identity review. Otherwise, the workflow turns one small drift into the new source of truth.

Use this handoff rule:

  1. Inspect the candidate last frame on all five identity dimensions.
  2. Confirm that the pose and crop are suitable for the next shot.
  3. Compare it with the original approved anchor, not only with the preceding second of video.
  4. Mark the frame approved handoff, record the source clip and frame time, and keep it immutable.
  5. If it fails, return to the approved anchor or repair a new bridge frame in the still-image layer.

Do not automatically extract the last frame from every accepted clip. A clip can be usable in an edit even when its final frame is a poor next-shot anchor—for example, the face is motion-blurred, a prop is occluded, or the character is cropped at the knees before a full-body shot.

The Artlist workflow recommends last-frame chaining and a review loop. The added production safeguard is to promote only a reviewed frame. “Previous” does not mean “approved.”

Where YingTu fits—and where it does not

The current English YingTu video workspace exposes controls for text-to-video, image-to-video, first-frame, paired first-and-last-frame, and reference-media configurations across visible Veo and Wan routes. Treat it as a bounded place to inspect route controls and run a harmless diagnostic after checking the current model and key requirements.

The visible control surface is not evidence that a character-consistency generation will succeed. No successful YingTu identity-preserving output was produced for this guide, so do not read the route names as a quality claim, privacy promise, or benchmark. Record the exact route, model, date, and input mapping you actually test.

If your task is choosing a broad route for an ordinary still rather than preserving one character, use the general image-to-video guide. If you still need to build a reusable still-image character pack, the consistent-character guide owns that earlier job.

Stop before uploading a real person or client asset

Use a synthetic or disposable test until the provider and project boundaries are clear. Stop and verify before uploading:

  • a real person whose informed consent and intended use are not documented;
  • a child or other vulnerable person;
  • a client character sheet, campaign, storyboard, logo pack, or unreleased product;
  • a licensed character or artwork without the required rights;
  • an ID, private location, medical context, financial context, or confidential background detail;
  • a reference when storage, retention, training use, deletion, visibility, subprocessors, or account ownership is unclear;
  • a likeness intended to deceive, impersonate, harass, or bypass approval.

For a real-person likeness, “the tool accepts the upload” is not rights clearance. Obtain the needed consent, releases, client approval, and provider-term review for the intended market and use. This workflow is not legal advice, and no visual match is worth crossing a consent or confidentiality boundary.

Three production examples

Mascot ad: start with the signature prop

A beverage mascot must turn, lift a bottle, and finish on a pack shot. Lock face, hair silhouette, body proportions, jacket, bottle label shape, and cel-shaded style. Prepare a three-quarter anchor, full-body view, and clean bottle reference in Nano Banana Pro.

Test the turn and lift as one short diagnostic. Review the bottle in the demanding middle frame; a correct face with a mutated bottle is a critical fail. If the final pack-shot composition is fixed, move to paired first/last frames only after the single-shot identity evidence is sound.

Short film: the occlusion is the hardest moment

A fictional detective walks behind a foreground pillar and reappears in profile. The opening portrait is not the hard part; reappearance is. Supply a profile reference, keep coat and hat construction locked, and use a short clip with a static camera first.

If the face changes after the pillar, reduce the walking motion and test the same occlusion. If the failure repeats with matching profile evidence, switch conditioning or redesign the shot. Do not accept a good first and last frame when the reappearance creates a different detective.

Approved avatar: the rights record is part of the asset

A company has written performer consent for a specific internal training series. Store the approval scope with the reference version, route, and review date. Lock face geometry, hair, body proportions, wardrobe, and realism level; allow expression and small gestures.

If the intended use expands to public advertising or a new market, stop and renew approval before generation. Technical consistency never enlarges the rights grant.

Copy this blank continuity record

hljs text
Project / shot:
Route, model, and date:
Approved reference version:
Rights / consent status:
Hardest required moment:

Locked identity traits:
Allowed changes:
Forbidden drift:

                    FIRST     HARDEST MOTION FRAME     LAST
Face:
Hair:
Body:
Outfit / props:
Visual language:

Motion verdict:
Camera verdict:
Timing verdict:
Ending-composition verdict:

Overall: PASS / REPAIR / SWITCH / STOP
Smallest reduced-variable retry:
Approved next-shot anchor or handoff frame:
Reviewer / approval date:

Keep accepted and rejected clips. A small evidence trail makes later drift diagnosable and prevents teams from rediscovering the same failed setup.

FAQ

Is one reference image enough to keep a character consistent?

It can be enough for one short shot that stays close to the reference angle and framing. It is weak evidence for a profile turn, full-body action, heavy occlusion, or a multi-shot sequence. Add only the views required by the hardest shot.

Should I use a first frame or first and last frames?

Use a single first frame when the shot begins from an approved composition and the ending is flexible. Use paired frames when the ending composition is essential and the route supports it. Paired endpoints still require review of the demanding middle frame.

When is reference-to-video the better route?

Use it when identity, outfit, prop, or style evidence needs to guide more than frame zero. Verify the selected route's current input mapping and limits; “reference” can mean different things across providers.

Can Nano Banana Pro generate the final video?

In this workflow, Nano Banana Pro prepares the still or reference pack. Motion comes from a separate video route. Keeping that boundary clear makes it easier to diagnose whether a failure belongs to the reference image, the shot request, or the video route.

Does repeating the same seed keep the same character?

Do not rely on it. A seed or repeated text prompt is not a substitute for approved visual evidence and frame-level acceptance. Use the reference method supported by the actual route and review the output.

Is it safe to use the previous clip's last frame as the next first frame?

Only after that exact frame passes all required identity checks and suits the next shot's framing. If it contains blur, occlusion, crop problems, or drift, return to the approved anchor or repair a bridge frame.

How many failed generations should I allow?

Set the limit before testing. A practical stop rule is to end a setup after the same critical identity dimension fails twice under reduced-variable conditions. Then switch route, strengthen the reference method, revise the shot, budget manual cleanup, or stop.

Can I use a real face or client character sheet?

Only after consent, rights, intended use, confidentiality, retention, and provider terms are acceptable for that asset. Use a harmless synthetic reference while any of those remain unclear.

Tags

Share this article

XTelegram