OpenShell 0.1 upgrades require a coordinated rebuild, with no mixed-version safety net
Operators must stop the fleet, preserve data and encryption material, migrate schema and credentials, then recreate sandboxes before validating clients and extensions.
NVIDIA’s OpenShell 0.1.0 upgrade is a coordinated migration, not a rolling update. Local 0.0.x installations cannot be upgraded in place, and deployed components from the two release lines cannot safely coexist because protocols, field names and time representations changed.
Sequence the cutover
Start by stopping every gateway replica and backing up the database. Export provider profiles before changing the installation; 0.1.0 removes built-in profiles, and aliases such as claude and gh become the canonical claude-code and github identifiers. Preserve the gateway’s key-encryption-key Secret as well: new provider credentials use the active credential driver, with encrypted database storage as the default.
Next, migrate gateway.toml to schema version 2. The file needs an explicit version, a singular compute-driver selector and driver settings under the new hierarchy. Run openshell-gateway config preflight before restarting. Helm users must put application configuration under gatewayConfig, while leaving TLS keys, passwords, client secrets, RBAC, Services, image settings and volumes in the chart-managed values and Secret interfaces.
Upgrade gateways, compute and credential drivers, supervisors, middleware, CLI clients and SDK clients from the same 0.1 release. Custom extensions must negotiate protocol 1.0 using PeerMetadata, advertise their family capability and regenerate bindings for renamed fields plus protobuf Timestamp and Duration values. Reusing 0.0.x generated bindings is explicitly unsupported.
Finally, delete and recreate every sandbox. Persisted 0.0.x sandbox boundaries and runtime descriptors are incompatible with 0.1.0. Reimport provider profiles at their original global or workspace scope and recreate credentials and refresh grants. Deployments using the Kubernetes Secrets credential driver get no credential migration; they require a fresh gateway and newly created provider credentials.
What breaks in a mixed deployment
A partial rollout can fail at several layers: old peers cannot decode renamed protocol fields or new time types; extensions may be rejected during capability negotiation; old sandboxes carry unverifiable runtime descriptors; and mixed gateways can corrupt assumptions around refresh records. Kubernetes workspaces can also fail sandbox bootstrap unless the matching workspace chart supplies Secret permissions.
User workflows change too. sandbox create --from no longer builds a local directory or Dockerfile, so users must build and push an image first. Provider attachment is explicit, managed inference routes are gone, unknown policy fields are rejected, and workspace selection can no longer rely on omission.
Validate before reopening service
Treat configuration preflight as the first gate, not the last. Then verify that every peer reports the same release and extension protocol, provider profiles resolve under the intended scope, credentials decrypt, and newly recreated sandboxes start with the expected runtime identity and network policy. Exercise one mutation twice with the same stable request ID, test deletion’s typed outcome, and reconnect a watch from its saved cursor. For interactive exec, close input and continue draining output until both the exit event and final RPC status arrive.
Only after those checks should users regenerate SDK clients, replace offset pagination with page tokens, and resume normal workloads. The safest rollback point is the stopped, backed-up 0.0.x system—not a mixed fleet.
sources
- Upgrade to NVIDIA OpenShell 0.1.0docs.nvidia.com
comments · 0