Control your token costs without slowing your engineers
PaletteAI Inference Launchpad is Spectro Cloud’s turnkey, local-first inference solution that runs modern open source models on either AMD or NVIDIA GPUs. Engineering teams keep working with the same IDEs, agents, and workflows but token spend doesn’t outrun engineering.


Join us at AMD’s Advancing AI from July 22-23 in San Francisco
Join us for a live demo at 515A or schedule a time with us through the event app for a personalized discussion.
Today
With PaletteAI Inference Launchpad
What it does
A single-cluster solution that runs a strong open model locally, on AMD or NVIDIA GPUs. Engineering teams get a familiar experience and platform teams get a real production service.
Inference Launchpad ships with:

Common inference tasks route to the local model. More complex work such as deep research or novel reasoning bursts out to a frontier model, determined by policies defined by the platform team.
How it works


At the foundation sits a single Kubernetes cluster on a validated AMD 325x or NVIDIA H200 GPU node, with a self-hosted open model served through an intelligent proxy that handles request routing, classification, and multi-tenant isolation. A local UI and an OpenAI-compatible API surface make the solution a drop-in target for the tools engineering teams already use.
Policies defined by the platform team determine where each request goes. Common inference tasks route to the local model and more complex requests such as deep research or novel reasoning burst out to frontier models.
Installation runs on the hardware that infrastructure teams already know how to procure, in connected or air-gapped configurations, and the solution integrates with the SSO, RBAC, and observability tooling that already surrounds every other production service in the data center.

Start fast, scale on the same platform
Inference Launchpad is the fastest way onto PaletteAI. A validated GPU node, boot the appliance, and it serves traffic in a day. That’s enough to prove the economics and help control token costs.
Scaling across teams, across sites, or offering inference as a governed internal service means moving to the PaletteAI platform.

Calculate your token cost savings
Engineering teams keep working with the same IDEs, agents, and workflows but token spend doesn’t outrun engineering.
