Control your token costs without slowing your engineers

PaletteAI Inference Launchpad is Spectro Cloud’s turnkey, local-first inference solution that runs  modern open source models on either AMD or NVIDIA GPUs. Engineering teams keep working with the same IDEs, agents, and workflows but token spend doesn’t outrun engineering.

PaletteAI logo

Join us at AMD’s Advancing AI from July 22-23 in San Francisco

Join us for a live demo at 515A or schedule a time with us through the event app for a personalized discussion.

Learn more

Today

Token costs grow exponentially with every prompt and every agent
Source code and prompts leave your perimeter
Building an inference stack in-house takes months
Cutting off developer access kills momentum

With PaletteAI Inference Launchpad

Turnkey solution enabling predictable token cost per developer
Everything runs inside the perimeter
Validated solution, serving traffic within a day
Same IDEs, same agents, same workflows!

What it does

A single-cluster solution that runs a strong open model locally, on AMD or NVIDIA GPUs. Engineering teams get a familiar experience and platform teams get a real production service.

Inference Launchpad ships with:

A pre-validated open model (GLM 5.2, DeepSeek v4 Pro, Gemma 4, Kimi K2.7), hot-swappable
An intelligent proxy for routing, quota, and multi-tenant isolation
A local UI and OpenAI-compatible API
Air-gap support and Day 2 operations built in

Common inference tasks route to the local model. More complex work such as deep research or novel reasoning bursts out to a frontier model, determined by policies defined by the platform team.

How it works

At the foundation sits a single Kubernetes cluster on a validated AMD 325x or NVIDIA H200 GPU node, with a self-hosted open model served through an intelligent proxy that handles request routing, classification, and multi-tenant isolation. A local UI and an OpenAI-compatible API surface make the solution a drop-in target for the tools engineering teams already use.

Policies defined by the platform team determine where each request goes. Common inference tasks route to the local model and more complex requests such as deep research or novel reasoning burst out to frontier models.

Installation runs on the hardware that infrastructure teams already know how to procure, in connected or air-gapped configurations, and the solution integrates with the SSO, RBAC, and observability tooling that already surrounds every other production service in the data center.

“PaletteAI Inference Launchpad is the only turnkey inference solution that runs on both AMD and NVIDIA on the same operational model, cutting inference costs by up to 70%, and backed by the same engineering discipline that runs modern infrastructure for governments, banks, and hospitals.”
Tenry Fu
CEO and Co-founder, Spectro Cloud

Start fast, scale on the same platform

Inference Launchpad is the fastest way onto PaletteAI. A validated GPU node, boot the appliance, and it serves traffic in a day. That’s enough to prove the economics and help control token costs.

Scaling across teams, across sites, or offering inference as a governed internal service means moving to the PaletteAI platform.

Learn more about PaletteAI

Calculate your token cost savings

Engineering teams keep working with the same IDEs, agents, and workflows but token spend doesn’t outrun engineering.