> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usestatemachines.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Recover interrupted management operations

> Use idempotency keys, explicit waits, and resource IDs when a call has an uncertain outcome.

A timeout does not prove that creation failed. The server may have accepted the request before the response was lost. Replaying the same create operation with the same idempotency key avoids creating a second environment.

## Preserve the operation identity

Choose an `idempotencyKey` before calling `environments.create()`. Save it with the exact create input if the caller must recover after a process restart. The SDK already reuses one key across its automatic retries.

Replay uncertain creation with the same key, input, workspace, and caller identity. An API key is part of that identity. Replacing it with a different key does not preserve the same create operation. A different input with that key fails with `idempotency_key_reused`. Use a new key only for a new operation.

```ts theme={null}
import {
  type CreateEnvironment,
  StateMachines,
  StateMachinesError,
} from '@usestatemachines/sdk';

const sm = new StateMachines();
const input: CreateEnvironment = {
  apps: ['salesforce'],
  name: 'Recoverable run',
};
const idempotencyKey = crypto.randomUUID();
console.log('Save this operation key for recovery:', idempotencyKey);

async function createAccepted() {
  try {
    return await sm.environments.create(input, { idempotencyKey, wait: false });
  } catch (error) {
    if (!(error instanceof StateMachinesError)) {
      throw error;
    }

    const unknownOutcome = error.status === null && error.code === null;
    const transientStatus =
      error.status === 502 || error.status === 503 || error.status === 504;

    if (!unknownOutcome && !transientStatus) {
      throw error;
    }

    return sm.environments.create(input, { idempotencyKey, wait: false });
  }
}

const env = await createAccepted();

try {
  const running = await sm.environments.wait(env.id, { status: 'running' });
  console.log(running.id, running.status);
} finally {
  await env.delete();
}
```

For restart recovery, persist the key and input outside the process before the first call. Printing the key as this demonstration does makes it visible, but is not a durable job record.

This example makes one explicit replay after the SDK exhausts its own retries. The SDK retries for up to 10 minutes first. If both calls fail without a known outcome, keep the printed operation key and the input. Recover that operation later. Do not replace the key just to get past the error.

A replay returns the same resource as it is now. If it was deleted, replay does not create a replacement. With the default wait, a replay of a deleted environment throws `UnexpectedStatusError`, or `EnvironmentFailedError` if it was deleted because it failed. Use `wait: false` to inspect a recovered resource before deciding the next action.

Snapshot creation accepts the same idempotency option. The key identifies the create operation, not the future native writes you make inside an environment.

## Separate acceptance from waiting

Use `{ wait: false }` when your code needs the environment ID before startup finishes. Put the subsequent wait inside the same `try` whose `finally` deletes the environment.

`WaitTimeoutError` means the SDK stopped waiting. The resource may still become ready. Read it again, wait again, or delete it when the task no longer needs it. Aborting a wait also does not undo an accepted operation.

## Classify the failure

Check errors with `instanceof`. `EnvironmentFailedError` carries the environment and its failure. `SnapshotFailedError` carries the failed snapshot. `UnexpectedStatusError` means the resource cannot reach the requested status without another action.

`StateMachinesError` provides `code`, `status`, `requestId`, `fields`, and `retryAfterMs`. Inspect the code and preserve the operation identity before deciding to replay. The [SDK error reference](/sdk-api/errors) describes each class and code.

The SDK retries transient management failures. Native app requests made with `credentials.fetch()` are never retried. After an uncertain Salesforce write, inspect the app state before deciding whether another write is safe.

## Bound waits without losing ownership

A wait timeout carries `resource`, the last successful resource read, which can be null. Keep the accepted ID separately even if no poll succeeds.

If you pass a `signal`, check its `aborted` state before you classify the error. Keep cleanup outside the aborted signal's scope.

For snapshot creation, also retain the accepted snapshot ID before waiting. A `saving` snapshot cannot be deleted. Either leave its source running and resume observation, or end the source and wait for the interrupted save to settle before deleting snapshot metadata.

A cleanup failure is a separate failure from the task. Preserve enough job state to retry deletion, and report the cleanup error without discarding the original task failure. A `finally` block cannot run after a process crash or a forced CI termination.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.