4:41 PM EDT
US Business Forum

Technology

How smaller, smarter models bring down the cost per token of high-volume AI

Business Insider Subscribe Account icon Account icon Search Business Strategy Economy Finance

Source: Business Insider5 min read
AI

Business Insider

Subscribe

Account icon

Account icon

Search

Business

Strategy

Economy

Finance

Retail

Advertising

Careers

Law

Media

Real Estate

Small Business

The Better Work Project

Personal Finance

Tech

Science

AI

Enterprise

Transportation

Startups

Innovation

Markets

News

Stocks

Indices

Commodities

Crypto

Currencies

ETFs

Lifestyle

Entertainment

Culture

Travel

Food

Health

Education

Parenting

Military & Defense

Politics

Reviews

Home

Kitchen

Style

Streaming

Pets

Tech

Deals

Gifts

Tickets

Video

Big Business

So Expensive

View From Above

Small Business

Authorized Account

Risky Business

Boot Camp

Still Standing

How Crime Works

Life Lessons

Games

Boxed

Pipeline

Knighthop

Tally

Stream Business Insider

Level Up

Community

Job Scout

Subscribe

My account

Log in

Newsletters

Saved Articles

US edition

Deutschland & Österreich

España

Japan

Polska

TW 全球中文版

Get the app

Sponsored content

How smaller, smarter models bring down the cost per token of high-volume AI

Sponsored by Dell AI Factory with Nvidia

BI Studios

2026-09-08T19:18:09.217Z

Read in app

Copy link

Email

Facebook

WhatsApp

X

LinkedIn

Bluesky

Threads

lighning bolt icon An icon in the shape of a lightning bolt.

Impact Link

Save Saved

An AI chatbot gives an answer to every question. An AI agent is far more autonomous and goes beyond simple answers; it makes plans, calls tools, checks its work, and repeats the loop until it's finished. Little of that thinking reaches the screen, but it has a cost in tokens.

Dr. Jon Krohn, fellow of machine learning practice at Lightning AI, a cloud-based provider of AI development tools, explains the math: cost per token, multiplied by how much background thinking the agent does, multiplied by how many data points it creates. "All these multiples mean token usage is exploding," Krohn says.

Total inference cost depends on the cost of processing each token, the number of tokens required per task, and the volume of tasks. Agentic workflows can increase all three, because agents plan, call tools, check results, and retry.

"If the model knows a little about everything, but nothing about your business, you'll get sensible-looking answers, but not the best answers," Krohn says.

An off-the-shelf model not fine-tuned to your business produces a generic agent. For a defined enterprise task, a smaller model trained on relevant data can be more efficient and accurate than a general-purpose model. What's more, it costs less per token and processes far fewer, because it isn't reasoning toward an answer it already knows.

Accuracy compounds the savings: A model that gets things right more often requires fewer people to check it. Small models keep GPUs busier, which means less wasteful downtime.

Krohn says the total benefit is not a few percentage points better — it's orders of magnitude better. Gartner predicts that organizations will use small, task-specific models three times as often as general-purpose LLMs by 2027.

The obstacle is the patchwork most organizations manage, characterized by underpowered devices, scattered data, and disconnected systems. Dell's full-stack AI approach with Nvidia fixes that with three components.

Deskside Agentic AI runs production-ready agents locally on enterprise hardware — where the token meter stops running. The AI Data Platform brings scattered enterprise data together into a foundation agents can actually use. And the AI Factory Foundation supplies validated architecture and tested deployment, so workloads don't need rebuilding every time they grow.

Lightning AI runs its own data centers and sells elastic capacity, rather than reserved cloud instances, the fixed blocks that cost money even when nothing is running. "We're fully integrated from the facility all the way to the application stack," says Frank Basso, Lightning's vice president of infrastructure.

The underlying hardware comes from the Dell AI Factory with Nvidia portfolio: liquid-cooled PowerEdge XE9780L and XE9785L servers built on Nvidia HGX B300, alongside Dell's newest rack-scale systems built on NVIDIA GB300 NVL72. Lightning was among the first providers to deploy the GB300 racks, packing upward of 25 megawatts and 10,000 GPUs into a single 10,000-square-foot data hall — a footprint that would have required roughly 200,000 square feet in a conventional air-cooled data center. The liquid cooling runs in a closed loop, like a car radiator, so the data center consumes far less water.

Dell PowerRack arrives from Dell's integration facilities already built and tested, and Dell field services certify them before handover, rather than Lightning assembling them on-site. For Krohn, that's what makes the economics real: The Dell AI Factory with Nvidia "allows economic benefits of training our own AI models to go from theoretical to actually be implemented," he says, because the cooling, GPUs, storage, and networking arrive as one kit.

Krohn's checklist for companies getting this right starts with understanding where efficiencies can be found. Model size and infrastructure efficiency can affect the cost of processing each token, while specialization can reduce the number of tokens, retries, and checks needed to produce a usable result. A winning strategy addresses both of these areas.

From there, start with evaluation, Krohn says. Know which workflow you want to change and how you'd measure whether a trained model improved it.

Next, treat training as a loop rather than an event, logging what goes wrong and feeding it into the next version. Finally, own the infrastructure.

With this sequence in place, Krohn and Basso agree, the payoff is both smarter models and exponential savings.

For high-volume inference, the potential savings come from combining smaller, task-specific models that use fewer tokens with an integrated infrastructure designed to keep computing resources productive and simplify deployment at scale.

Discover how Dell and Nvidia can help turn smaller, smarter AI models into lower-cost, production-ready agents.

This sponsored post was created by BI Studios with Dell AI Factory with Nvidia.

HOME

Subscribe

Legal & Privacy

Terms of Service Terms of Sale Privacy Policy Accessibility Code of Ethics Policy Reprints & Permissions Disclaimer Advertising Policies Conflict of Interest Policy Commerce Policy Coupons Privacy Policy Coupons Terms

Company

About Us Careers Advertise With Us Contact Us News Tips Company News Awards Masthead

Other

Games Sitemap Stock quotes by finanzen.net Corrections Community Guidelines AI Use BI Live Business Insider App

International Editions

AT DE ES JP PL TW

Copyright © 2026 Insider Inc. All rights reserved. Registration on or use of this site constitutes acceptance of our Terms of Service and Privacy Policy.

Jump to