TrySora guide

Sora AI text to video: write one controllable shot

A useful text-to-video prompt is a compact production brief. It defines what appears, what moves, how the camera behaves, and what must remain stable without hiding the actual model selected in the workspace.

Use the live provider and model label as the technical boundary. The tutorial works across supported video models; it does not promise that every task uses an OpenAI Sora model.

External product and API facts reviewed August 27, 2026.

Build the prompt in six decisions

Write the subject and environment first, followed by framing, subject action, one camera move, atmosphere or sound, and constraints. Keeping the order stable makes a failed result easier to diagnose.

Example: A small electric coupe waits on a rain-dark street at blue hour, low three-quarter medium shot. Steam drifts behind the stationary car. The camera tracks slowly left to right. Soft traffic ambience, physically plausible reflections, no cut, no text, preserve the wheel and body geometry.

  • Subject and setting
  • Framing and viewpoint
  • One observable action
  • One camera movement
  • Light, atmosphere, and optional sound direction
  • Preservation and exclusion constraints

Choose text or image input deliberately

Use text-to-video while the art direction is open. Use image-to-video after a character, product, wardrobe, or composition is approved. With an image reference, spend fewer words redesigning the frame and more words on motion and preservation.

Before submitting, read the available model, ratio, duration, and credit estimate in the live workspace. A tutorial cannot establish settings that the selected provider does not expose.

Review one failure category at a time

Watch the result once for subject action, once for camera behavior, and once for geometry and background continuity. Inspect faces, hands, text, logos, product shape, and the final edit point. Change one instruction between comparable attempts so you can tell what improved the shot.

  • Wrong motion: revise the action phrase.
  • Camera drift: simplify or lock the camera.
  • Identity drift: strengthen preservation or start from an approved frame.
  • Busy result: remove secondary actions and extra scene changes.

Frequently asked questions

How long should a Sora AI video prompt be?

Long enough to define one shot, action, camera move, atmosphere, and essential constraints. Remove language that does not change an observable decision.

Should I put several scenes in one prompt?

Start with one shot. Several locations, cuts, and actions make it harder to identify why a result failed and harder to reuse a good prompt.

Does this tutorial guarantee the Sora 2 model?

No. It is a model-aware workflow. The exact provider and model shown before submission determine what actually runs.

When should I use an image reference?

Use one when appearance or composition is already approved and motion is the remaining decision. Upload only material you have permission to use.

Continue with a related guide