Description
Deploy the current workspace, or activate an existing release. With no positional argument, this is the normal deployment path:- reads the model and release from
reactor.yaml, - registers the model when it does not exist yet,
- publishes the release if it does not already have an image,
- applies the current model configuration, and
- activates the release using
deployment.yamlwhen present.
reactor init, the usual flow is one command:
<model>:<release> skips reactor.yaml publication and
configuration updates, then activates an already-published release. A local
deployment.yaml still controls its regional instance plan when present.
Use that form for rollbacks and CI flows that publish separately; it cannot be
combined with --source.
Use this command to:
- Activate a freshly published release when it is ready for traffic.
- Roll back to an older release if the current one has issues.
reactor weights upload.
When the workspace deploy publishes the release, the publish step reads
runtime.weights_path from reactor.yaml the way reactor model publish does.
--weights <path> (or an hf://org/repo[@rev] reference) names a different
source for this release; --no-weights skips the step, for a machine that
holds no weights.
Build flags:
When the workspace deploy builds the release, --build-env and
--build-secret pass to that build with the same grammar as
reactor model publish. An already-published release skips the
build, so they have no effect.
Workspace-aware:
When run from a workspace (any directory under one containing a
reactor.yaml, stopping at the enclosing repository root) with no
positional argument, the model name, release tag, mutable model configuration,
and image build are taken from the workspace.
An explicit <model>:<release> activates that existing release without
rebuilding it or applying the workspace spec. Its <model> part must
match the workspace model.name or the command is rejected.
Strategy:
By default a deploy replaces instances one at a time, so the model keeps
serving throughout. --strategy replace-all signals the whole old generation
at once: an idle fleet turns over in seconds, and a busy one runs at reduced
capacity until the fresh instances finish loading. Under both, an instance
still serving a session keeps serving it until the session ends or the
instance’s grace period runs out, which is up to five minutes in production.
Immediate (emergency cutover):
--immediate cuts that tail. After signalling the rollout, the platform
disconnects the model’s live sessions so the new release takes over in
seconds instead of minutes. Reach for it when the running release must stop
serving now (a bad version, a safety problem), not for routine deploys.
It implies --strategy replace-all (spelling the strategy out is fine;
any other strategy is an error). It terminates whatever sessions are live
when the sweep runs, including ones that started during the deploy. Two
overlapping deploys, or a retry, each disconnect the then-live set.
Clients are told the model was redeployed where their session’s release
supports it; on older releases they see a plain disconnect. Their sessions end
either way.
--yes skips the confirmation prompt for scripts and CI.
Examples:
Options
Global options
See also
- reactor model - Model management commands