What's a token worth to a bank? AI ROI and token economics, live from New York
If you run AI in a bank, insurer or trading firm right now, you're fielding two questions at once. The board wants to know why more of your AI work isn't in production yet. Finance wants to know why the AI that is in production costs so much. Answering either one means getting specific about token economics, and that's exactly what we went to New York to talk about.
On September 24, Ignite Community brought banking, insurance, capital markets and fintech leaders together at Current, Chelsea Piers for AI for Financial Services, an executive event hosted by Supermicro and AMD. Our CEO and co-founder Tenry Fu joined the final panel of the day, "Optimizing AI ROI + token economics," alongside:
- Shipali Jangra, Director of Digital Products, American Express
- Steve Hou, Head of Research, Silicon Data
- Virju Patel, Founder and CEO, Data Maverick (moderator)
Virju opened by framing enterprise AI as three doors: frontier models consumed as external intelligence from the OpenAIs and Geminis of the world, open models you build and run privately in your own data center, and the hybrid of both. The panel's job was to work out what walking through each door costs, and who should be keeping the receipts. The full 25-minute conversation is now on YouTube, and it's worth your time. Here's where it went.
Ninety percent of pilots are still pilots
Shipali brought the customer and product view, fresh from 2027 planning cycles where "how do you justify the ROI?" is being asked out loud. Proofs of concept aren't the problem; there are plenty of excellent ones. Scaling them is where things stick. She shared a data point from an internal committee: more than 90% of use cases still sit in POC and never truly scale to production. Her diagnosis for the regulated world: treating governance as an afterthought doesn't work. Privacy and governance have to be baked in from the start, or the pilot dies the moment it touches a production environment with a regulator watching.
Tenry's take is that production is where the overlooked work lives. "It's not just a one-time deployment," he argued. It's how you handle updates, scale, support and security, atomically, across a stack that now mixes GPUs, high-speed networking, inference engines and models. And since the vast majority of enterprises will never train frontier models, their AI spend concentrates on inference, which makes token costs, inference performance and routing governance the core economic questions.
Steve was the optimist of the three, on classic economist grounds: revealed preference. People keep using these tools without anyone pointing a gun at their heads, so they're deriving value. The gap is in finishing. Vibe coding a dashboard is easy; closing, taking it through the finish line, is hard, and most people aren't good at it yet. His forcing function: the labor market. People who can carry AI work to done will get paid more.
Tokens are a terrible measure of intelligence
Virju's sharpest question: what's the right unit of measurement, a token, a GPU hour, or a finished task?
Shipali rejected the menu. At American Express the question is whether the decision the model made is trustworthy and auditable, because with GDPR and EU and UK regulation in play, an untrustworthy decision delivers no business outcome no matter what you spent on it.
Tenry's answer: whatever you measure, it starts with metering. You need full visibility into who is spending tokens, how many, on which model, before ROI is even a conversation. And because the gap between frontier and open-weight models keeps shrinking, most enterprises will run several models at once, so no single vendor's usage dashboard can give you the picture. Metering has to sit in a layer you control, across all of them.
Steve went further: "Token is a terrible measure of intelligence over a short period of time." Tokens are deflating; each one buys more capability every month, faster than you can track. His preferred measure is defiantly traditional: does it show up in quarterly earnings? Can you find more high-value customers and serve them better? By that standard, he was honest: "the economics hasn't quite shown out yet." Shipali backed that up with a story from their Copilot rollout. They assumed their engineers would just figure it out; the survey came back with a significant number saying they didn't know how to use it. Her assessment of how far in we are: "Are we even at 0.1% of what this model or tool can do for us? The answer is no, at least for us."
Nobody on stage claimed to have solved ROI measurement. The honest position was that the unit economics are still being worked out while the bills are very real today. We've written before about why cheaper tokens keep producing bigger bills; this was that argument playing out among practitioners.
Rent or buy? The New York apartment question
Steve compared the compute decision to renting versus buying an apartment in New York, then described a market that's stopped behaving. A couple of years ago an AI startup could reserve a year of compute; now it's three years in one go with 33-40% cash down. "They're raising VC equity money to buy compute," he said. "That is honestly quite insane. This thing should be debt financed." Tenry added the inversion that surprises everyone: on the neoclouds, a long-term commitment now carries a premium over the on-demand price, because guaranteed GPU access is the scarce commodity. All of it traces back to supply chain constraints, and it's one reason he's seeing data center repatriation pick up.
The panel's buy-versus-rent logic ended up refreshingly simple. Buy, and your cluster becomes what Tenry called "a fixed throughput token factory" with roughly a three-year life: great economics if you can keep it saturated, depreciation on your books if you can't. Rent if your token spend is still unpredictable. Steve added the caveat that buying means taking up a side business of running a data center, with the people and expertise that implies, though at current rental prices the payback period is getting short enough that plenty of firms are buying anyway.
Where the panel thinks this lands next year
Asked for a one-year outlook, Steve flagged concentration as the risk to watch: he cited stats from Apollo he'd seen that morning suggesting around 10% of companies account for some 95% of compute demand for OpenAI tokens. If enterprise diversification doesn't catch up fast enough, there's an air pocket. He hopes it changes in a meaningful way; the moderator heard "medieval," which by some readings would also fit.
Shipali doesn't expect the world to change drastically in a year. The win she wants is fit: marketing's creative campaign doesn't need the most advanced model on the market, while credit fraud analysis that hinges on anomaly detection might. Clarity about which use case deserves which tool would, in her words, be a big win for large organizations by itself.
Tenry closed with the view we hold at Spectro Cloud: this all ends in hybrid AI. Just as the cloud era settled into multicloud, enterprises will run AI everywhere the business needs it, from desktop to edge to data center, colo, neoclouds, sovereign clouds, hyperscalers and frontier labs, choosing the right silicon for the right workload and the right model for the right inference, with one control plane to govern the lot. That's the design brief behind PaletteAI Inference Launchpad: private inference with intelligent routing across local and frontier models, and metering so you can see where the tokens go.
Watch the full conversation
The whole panel is on our YouTube channel. If the economics questions hit close to home, the Inference Launchpad launch announcement is a good next read, and if you'd like to talk through what private inference could look like in your environment, we'd love to hear from you.
Thanks to Ignite Community, Supermicro and AMD for a well-run event, and to Shipali, Steve and Virju for a conversation that didn't pull punches.

