Business-critical optimization workloads often require continuous availability, even in the face of unexpected cloud service disruptions. To achieve this, Gurobi Instant Cloud allows you to deploy multiple pools, each in a different cloud region, so that your applications can automatically fail over if one region becomes temporarily unavailable.
This article explains:
- Why regional redundancy is important
- How to structure multiple pools across regions
- How to implement automatic failover in your client application
- Optional advanced techniques using the Gurobi Cloud REST API
Why High Availability Matters
Cloud providers occasionally experience regional disruptions, network partitioning, or API rate issues that can affect resource availability. By configuring multiple Instant Cloud pools across regions, you can ensure that your optimization jobs continue to run smoothly, even if one region is impaired.
Each Instant Cloud Pool encapsulates:
- Machine type (e.g.,
m6i.large) - Cloud region (e.g.,
us-east-1,us-west-2,eu-west-1) - Pool settings such as idle shutdown, job timeouts, and scaling behavior
When a client connects to a pool, Gurobi Cloud automatically launches or reuses compute machines in that pool.
Creating a Pool in Gurobi Instant Cloud
Before you can configure failover, you need to create the individual Instant Cloud pools that your clients will connect to. A pool defines a reusable configuration of compute machines (region, size, and other settings) that can be shared across users in your organization.
You can create pools in one of two ways:
1. Using Gurobi Cloud Manager
Create a Pool in Gurobi Instant Cloud
- Sign in to Gurobi Cloud Manager.
- In the Pools section, click “Create Pool”.
- Choose:
-
Region (e.g.,
us-east-1,us-west-2, oreu-west-1) -
Machine type (e.g.,
m6i.large) - Idle shutdown and Job timeout parameters
-
Region (e.g.,
- Click Save to create the pool.
- It will appear in your list with a unique Pool Name (e.g.,
prod-us-east-1). - (Optional) Repeat the process to create one or more secondary pools in other regions.
2. Using the REST API (Advanced)
Automate Pool Creation via the REST API
You can also automate pool creation using the Instant Cloud REST API.
The endpoint to create a pool is:
POST https://cloud.gurobi.com/api/v2/pools
Example JSON body:
{
"name": "string",
"licenseId": "string",
"licenseType": "light compute server",
"description": "string",
"provider": "aws",
"region": "us-east-1",
"machineType": "c5.4xlarge",
"GRBVersion": "latest",
"idleShutdown": 0,
"idleJobTimeout": 0,
"jobLimit": 0,
"nbComputeServers": 0,
"nbDistributedWorkers": 0,
"jobHistory": true
}
Reference: Interactive Swagger Console
Once your pools exist, you can connect to them from any Gurobi client using their Pool Name, along with your Cloud Access ID and Secret Key.
A Standard Multi-Region Pool Strategy
| Purpose | Pool Name | Region | Typical State |
|---|---|---|---|
| Primary | prod-us-east-1 |
us-east-1 |
Active |
| Secondary | prod-us-west-2 |
us-west-2 |
Standby |
Both pools should:
- Use the same configuration (machine type, job queue, version)
- Share the same Cloud Access ID and Secret Key
- Differ only by their cloud region
Your application can connect to either pool simply by changing the CloudPool parameter in your license file.
gurobi-us-east.lic
# Gurobi Instant Cloud license (US East) # Save this file as gurobi.lic in a default location like your home directory # or set the GRB_LICENSE_FILE environment variable to the file's full path CLOUDACCESSID=XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX CLOUDSECRETKEY=YYYYYYYY-YYYY-YYYY-YYYY-YYYYYYYYYYYY CLOUDPOOL=prod-us-east-1
gurobi-us-west.lic
# Gurobi Instant Cloud license (US West)
# Save this file as gurobi.lic in a default location like your home directory
# or set the GRB_LICENSE_FILE environment variable to the file's full path CLOUDACCESSID=XXXXXXXX-XXXX-XXXX-XXXX-XXXXXXXXXXXX CLOUDSECRETKEY=YYYYYYYY-YYYY-YYYY-YYYY-YYYYYYYYYYYY CLOUDPOOL=prod-us-west-2
[Simple Client Side Approach]
Automating Failover via Client Environment Settings
You can automate failover directly in your client application; no REST calls required. Your code attempts to start a Gurobi environment in the primary pool; if startup fails or takes longer than a configured timeout, it automatically falls back to the next pool in your list (e.g., a different region).
Key parameters (set on an empty Env at runtime or via license):
CloudAccessIDCloudSecretKey-
CloudPool(the pool name, which determines the region)
This approach:
Automatically retries through multiple pools until one connects
Needs only environment parameters; no API setup
Works across any number of backup regions in priority order
Enforces a per-pool startup timeout to avoid long hangs
The example below implements this logic: if Env.start() exceeds the timeout, the client switches to the next pool automatically. Adjust the timeout and pool order to match your SLA and regional preferences.
Python Implementation
import threading
import queue
import gurobipy as gp
from gurobipy import GurobiError
# --- configure ---
ACCESS_ID = "<your-access-id>"
SECRET_KEY = "<your-secret-key>"
POOLS = ["prod-us-east-1", "prod-us-west-2", "prod-eu-west-1"] # priority order
ENV_START_TIMEOUT_S = 60 # how long to wait for env.start() per pool
# -----------------
def _start_env_worker(pool_name: str, out: queue.Queue):
"""Run env.start() off-thread so we can enforce a timeout."""
try:
env = gp.Env(empty=True)
env.setParam("CloudAccessID", ACCESS_ID)
env.setParam("CloudSecretKey", SECRET_KEY)
env.setParam("CloudPool", pool_name)
# Optional: reduce client verbosity while failing over
# env.setParam("OutputFlag", 0)
env.start()
out.put(("ok", env))
except Exception as e:
out.put(("err", e))
def start_env_with_timeout(pool_name: str, timeout_s: int):
q = queue.Queue(maxsize=1)
t = threading.Thread(target=_start_env_worker, args=(pool_name, q), daemon=True)
t.start()
try:
status, payload = q.get(timeout=timeout_s)
except queue.Empty:
raise TimeoutError(
f"env.start() timed out for pool '{pool_name}' after {timeout_s}s"
)
if status == "ok":
return payload
raise payload if isinstance(payload, Exception) else RuntimeError(payload)
def solve_with_failover(build_model_fn):
"""
Try each pool:
1) start environment with timeout
2) build model bound to that env
3) optimize
Fall back on any error or startup timeout.
"""
last_err = None
for pool in POOLS:
print(f"Trying pool: {pool}")
try:
env = start_env_with_timeout(pool, ENV_START_TIMEOUT_S)
except Exception as e:
print(f" env.start() failed/timed out on {pool}: {e}")
last_err = e
continue
try:
model = build_model_fn(
env
) # e.g., lambda env: gp.read("model.mps", env=env)
model.optimize()
print(f"Success on pool: {pool}")
return model, pool
except GurobiError as e:
print(f" Optimize failed on {pool}: {e}")
last_err = e
# try next pool
except Exception as e:
print(f" Unexpected error on {pool}: {e}")
last_err = e
# try next pool
raise RuntimeError(f"All pools failed. Last error: {last_err}")
# Example usage:
# model, used_pool = solve_with_failover(lambda env: gp.read("model.mps", env=env))
# print("Solved using pool:", used_pool)
[Advanced/Optional]
Using the Gurobi Cloud REST API to Check or Pre-Warm a Backup Pool
In addition to client-based failover, you can use the Gurobi Instant Cloud REST API (v2) to proactively monitor and prepare your backup pools. (See the Interactive Swagger Console)
This is especially useful if your workflow requires instant availability or if you want to automate health checks as part of your uptime strategy.
The REST API allows you to:
List your existing pools
Check whether a specific pool has active machines
Start (or “pre-warm”) a backup pool in advance of use
Optionally, wait until a machine is fully ready before connecting
⚠️ Important:
Starting a machine through the API activates a running environment, which incurs billing time. Gurobi Cloud currently bills a minimum of 30-minutes per environment startup, even if the machine is used for a shorter period. For this reason, pre-warming should be used sparingly, only when uptime requirements justify the additional cost.
This example demonstrates how to use the Gurobi Instant Cloud REST API (v2) to check the status of a backup pool and, if necessary, start a machine to pre-warm it before your client applications attempt to connect.
Authenticate: Use your Cloud Access ID and Secret Key (the same credentials that Gurobi clients use).
List your pools: Retrieve all Instant Cloud pools and find the one matching your backup region by name.
Check pool activity: Query the pool’s
/machinesendpoint to see if any compute machines are currently running.Start a machine if needed: If the pool has no active machines, send a
POSTrequest to launch one.Wait until ready: Poll the pool periodically until at least one machine is running (or “ready”), or until a timeout expires.
Connect normally: Once the pool reports an available machine, your standard Gurobi client or failover logic can connect immediately with minimal startup latency.
Python Implementation
import time
import requests
from requests.auth import AuthBase
# ---- Custom auth (adds v2 headers while keeping auth=auth call sites) ----
class GurobiCloudAuth(AuthBase):
def __init__(self, access_id: str, secret_key: str):
self.access_id = access_id.strip()
self.secret_key = secret_key.strip()
def __call__(self, r: requests.PreparedRequest) - requests.PreparedRequest:
r.headers["X-GUROBI-ACCESS-ID"] = self.access_id
r.headers["X-GUROBI-SECRET-KEY"] = self.secret_key
r.headers.setdefault("Accept", "application/json")
r.headers.setdefault("Content-Type", "application/json")
return r
# --- configure ---
ACCESS_ID = ""
SECRET_KEY = ""
BASE_URL = "https://cloud.gurobi.com/api/v2"
POOL_NAME = "prod-us-west-2" # backup pool
DESIRED_COUNT = 1 # how many machines to warm-start
READY_TIMEOUT_S = 180
POLL_EVERY_S = 5
# -----------------
auth = GurobiCloudAuth(ACCESS_ID, SECRET_KEY)
READY_STATES = {"idle", "running", "busy", "ready"} # "ready enough" states
def machine_state(m: dict) - str:
# Prefer "state", fallback to "status"; lowercased
return (m.get("state") or m.get("status") or "").lower()
def get_machines(pool_id: str) - list[dict]:
r = requests.get(f"{BASE_URL}/pools/{pool_id}/machines", auth=auth, timeout=30)
r.raise_for_status()
raw = r.json()
machines = raw["computeServers"]
return machines
def any_ready(machines: list[dict]) - bool:
return any(machine_state(m) in READY_STATES for m in machines)
def count_ready(machines: list[dict]) - int:
return sum(1 for m in machines if machine_state(m) in READY_STATES)
# 1) Locate the pool
resp = requests.get(f"{BASE_URL}/pools", auth=auth, timeout=30)
resp.raise_for_status()
pools = resp.json()
pool = next((p for p in pools if p.get("name") == POOL_NAME), None)
if not pool:
raise RuntimeError(
f"Pool {POOL_NAME} not found. Available: {[p['name'] for p in pools]}"
)
pool_id = pool["id"]
# 2) Check existing machines and readiness
machines = get_machines(pool_id)
if any_ready(machines):
print(
f"Pool {POOL_NAME} already ready "
f"({count_ready(machines)}/{len(machines)} machine(s) in ready state)"
)
else:
print(f"No ready machines in {POOL_NAME}. Launching {DESIRED_COUNT}...")
launch = requests.post(
f"{BASE_URL}/pools/{pool_id}/machines",
auth=auth,
json={"machineCount": DESIRED_COUNT},
timeout=30,
)
launch.raise_for_status()
# 3) Poll until at least one machine is ready
deadline = time.time() + READY_TIMEOUT_S
while time.time() = 1:
print(f"Pool {POOL_NAME} is ready ({ready}/{total} ready).")
break
time.sleep(POLL_EVERY_S)
else:
raise TimeoutError(f"Pool {POOL_NAME} did not become ready in time.")
# Continue with model building and execution with launched pool...We recommend that most customers consider the following best practices:
Poll carefully: Add a timeout and short sleep interval between checks (5–10 seconds).
Check machine state: When available, verify the machine’s
"status"field (e.g.,"running"or"ready") instead of just presence.Avoid unnecessary launches: Don’t pre-warm every job; only for mission-critical or latency-sensitive workloads.
Use environment variables for credentials; never hardcode your Access ID or Secret Key.
Integrate with monitoring: You can schedule this check as part of a cron job, AWS Lambda, or Azure Function to keep your secondary region pre-warmed during business hours.
Summary
Failover in Gurobi Instant Cloud is straightforward and powerful. By defining multiple pools in separate regions and using a short retry loop in your client application, you can ensure that your optimization workloads remain reliable, even when a specific region is unavailable.
We recommend that most customers consider the following best practices:
- Maintain at least two pools in distinct regions
- Keep pool configurations identical except for region
- Use short connection timeouts and retry logic for fast failover
- Test failover regularly by forcing connections to the secondary region
- Monitor usage and uptime through Gurobi Cloud Manager
For additional guidance on multi-region deployment, contact your Gurobi Technical Account Manager (TAM) or visit the Gurobi Cloud Documentation.