Skip to main content
Visko Orbis Dynamic renders continuous takes. Each prompt is one unbroken shot; the model streams it in 33-frame chunks and applies any new prompt at the next chunk boundary, so a mid-run set_prompt morphs the scene instead of cutting to a fresh one. The rules below come from the model team’s own prompting discipline, applied to Reactor’s wire surface. Two working rules govern everything on this page:
  1. Build the world once, then change it one clear step at a time. Your first prompt does the heavy lifting; every later prompt describes only what visibly changes.
  2. The model renders physical nouns and verbs, never adjectives of intent. Write what the camera sees, not what you want it to feel.

Initial prompt

The formula is always the same: WHO does WHAT at WHERE, plus camera control. The first prompt stakes the whole scene — the model builds on it, so missing pieces are what the model fills in however it likes. Aim for 1–3 sentences and fewer than ~100 words. There is no hard input limit; the model compresses what you send. Spend the budget on the details that matter, because the model gives the most space to what you give the most space.

Subject (the WHO)

Commit to one main subject, described with concrete appearance: age range, hair, skin tone, clothing, colors, materials, distinctive features. Fewer than four characters per scene.
A person streaming.
An influencer in her early 20s with straight, shoulder-length honey-brown hair parted down the middle, fair skin, natural makeup, and a warm, toothy smile, wearing an off-white chunky-knit crewneck sweater.

Background (the WHERE)

Control the elements that matter; let the model invent the rest. Useful elements to describe can be the environment, key background objects, lighting conditions, textures, atmosphere, and spatial layout. Don’t describe everything, only describe the parts that anchor the composition.
A cozy streaming studio.
A warm, aesthetic indoor streaming studio. In the foreground, a smooth white marble surface holds a small potted plant, a lit amber candle, stacked hardcover design books, and a small vase with eucalyptus. An illuminated black ring light stands on the right. The softly blurred background features three framed coastal and botanical art prints on a beige wall, a warm table lamp on a wooden cabinet, a black metal shelf with plants, and a gray sofa by a sunlit window.

Camera (the shot)

You don’t need to specify every setting, but name the controls that matter. The model takes instructions across four axes (any mixture works, listed as plain words you can type into the prompt): A complete camera line: "Medium shot, overhead, static camera, deep depth of field."

Rendering style

Default is photorealism. Only name a rendering style when you want one: cartoon style, anime style, etc. If you need very specific stylization (e.g. painterly, comic-book), prefer image-to-video over text: style guidance from a reference frame is far stronger than from words.

Image-to-video (I2V)

The image pins the first chunk; every later chunk inherits it. Start with text-to-video (T2V), then refine with I2V. T2V is the fastest way to learn whether the concept and prompt work at all; once the scene behaves, switch to a reference frame for tighter visual control. If your reference image is far outside the model’s familiar domain, generation may drift toward what it handles more naturally. Ensure the image’s aspect ratio is 16:9 landscape, and keep imagery close to photorealism even when the final look is stylized. Complete initial prompt example, combining WHO + WHAT + WHERE + camera:
A young woman in an off-white knit sweater sits at a wooden café table beside a large window, smiling to the camera. Warm morning sunlight enters through the window. Medium shot, eye-level, static camera, shallow depth of field.

Following prompts

Once generation is running, you do not need to describe the world again. The model treats every established element as still present and preserves it. The one question to answer is What do you want to visibly happen next?

One action per prompt

Use a single main action. If the action is complex, break it into steps across multiple prompts. Small actions are far more controllable than compound ones.
The man stands up from the bench, starts running along the trail, runs toward the park exit as the city street becomes visible beyond the trees, then leaves the park and continues onto the sidewalk.
A man sits on a bench inside a park. The man stands up from the bench and starts running along a trail. The man runs toward the park exit, with the city street becoming visible beyond the trees. He runs out of the park and onto the sidewalk.

State transitions, not restarts

Think of the world as a sequence of continuous state transitions. Each new prompt describes how the current state becomes the next state: the scene you see, plus one clear change.
  • From “a man sits on a bench”: The man stands up from the bench and starts running along a trail.
  • Continuing forward: The man continues running along the trail toward the lake.
That’s a natural progression: sitting → standing → running.

Scene transitions

Abrupt location swaps (“cut to: the sidewalk”) read as cuts. Bridge them through traversal: let the camera see the next location emerging before the move commits.
He is on the sidewalk. (From inside the park, teleport.)
The man runs toward the park exit, with the city street becoming visible beyond the trees.He runs out of the park and onto the sidewalk.

Introducing new subjects

Bring a new subject in through an action, not as background description:
A cat playing with the boy in the grass.
A cat enters the scene and begins playing with the boy.A sea turtle slowly swims into view beside the diver.
Once a subject has entered, refer to it naturally ("The turtle swims alongside the diver"). By default, every subject stays in the scene. To remove one, say it explicitly (“…as the cat walks out of frame and leaves the room”).

Changing established elements

If you want an existing element to change, say it explicitly, otherwise the model preserves the scene.
The woman is wearing a red jacket now. Sunny. Rain.
Change the woman’s black jacket to a red jacket.The weather changes from sunny to rainy.
Without that kind of instruction, subjects, backgrounds, objects, and camera all stay as they were.

The world moves at its own pace

Changes are not instant. After a new prompt, the model transitions the world over several chunks, usually 2–4 seconds. Give changes enough space to land before layering on another prompt.

Anti-patterns

Redescribing the world on every following prompt

The one hard rule for following prompts: do not restate the world. The model preserves what you do not change. Restating WHO or WHERE burns part of your input budget and can read as a scene rebuild. Describe how concrete details in the scene evolve and let the model carry forward the rest.

Compound actions in one prompt

He stands, runs, exits, turns left, catches the bus.
He stands up from the bench.He starts running along the trail.He turns left at the fork.
Compound actions read as four state transitions pretending to be one; the model collapses them to a blurry morph. Break them into adjacent prompts and let the world land between each.

Abrupt scene cuts

Now we are on the beach.
The man runs toward the park exit, the beach visible beyond the trees.He runs onto the sand.
A teleport reads as a cut. Bridge locations through traversal: let the camera see the next location before the move commits.

Terminally terse prompts

A girl at a café.
A young woman in an off-white knit sweater sits at a wooden café table beside a large window, smiling to the camera. Warm morning sunlight enters through the window. Medium shot, eye-level, static camera, shallow depth of field.
With no world to preserve, the model invents a new café every chunk and the frame drifts. Always WHO + WHAT + WHERE on the first prompt.

Adjectives of intent

Make it look cinematic. Make it more dynamic.
Slow push forward at eye level as the camera closes in on the glass.Backlit, rim light tracing the subject’s shoulders, the room falling to shadow around her.
Adjectives like “cinematic” or “dynamic” are wishes about the model’s own output, not descriptions of the world. The model renders nouns and verbs. Translate every adjective of intent into what the camera would actually see.

Negation

An empty street with no people.
A narrow street, the sidewalk empty.
The model renders the noun in no people; “street without people” places people. When you need something absent, name the positive state you want instead. If a noun has a known failure mode, name its opposite.

Setting an audio prompt from the scene description

Feeding the visual prompt into set_audio_prompt makes the audio measurably worse than leaving it unset. When unset, the audio model generates sound from the picture alone. If you do author an audio prompt, write one sentence about what it sounds like (instruments, voices, materials, ambience), never what it looks like. See the schema set_audio_prompt.

Layering steering faster than chunks land

Changes live on chunk boundaries plus one queued output chunk, which is around 2–4 seconds end-to-end before a new prompt is visible. Sending a new action every 500 ms just queues competing morphs. Give each prompt a few seconds to land before layering the next.

Forgetting to re-drive start after generation_complete

When a run reaches max_chunks (up to 229, ~7 min) the model emits generation_complete and returns to WAITING. It does not auto-restart. Send start again (same conditions) or reset first. The tutorial’s NowPlaying surfaces a “Run finished” CTA for exactly this reason.

Runtime caveats

  • The first chunk emits 0 frames (frames_emitted: 0). The stream needs a warm-up chunk before real pixels start flowing; first picture lands by chunk 2. Hold a “priming” overlay — don’t treat this as an error.
  • set_resolution and set_seed apply at the NEXT start and survive reset. The prompt and image do not survive reset. Reading these wrong makes your UI disagree with the model.
  • Non-16:9 reference images squash to 832×480 with no crop. Crop to 16:9 before upload, or accept the distortion.

See also