Not Every AI Feature Needs to Be an API Call

The next question for AI products is not which model to call, but whether the work needs to leave the device at all.

Published

October 4, 2026

Categories

AIInfrastructure

Written By

Demo Author

For the last few years, the default architecture for adding AI to an application has basically been: the user does something, you send it to somebody else's server, somebody else's GPU thinks about it, and you send the answer back. Done.

And honestly, that makes sense. Frontier models are massive, and running them yourself is expensive. APIs made it possible for somebody like me to build things that would've needed an entire ML team not that long ago. I love APIs. I use them constantly.

Cloud AI shouldn't be automatic

But I think we've gotten so used to cloud AI that sometimes we treat it as the automatic answer, and I don't think it should be. Not anymore. Models are getting smaller, hardware is getting better, and tooling for running inference locally is becoming much less painful.

We're getting to an interesting point where the architectural question isn't just which model should I call? It's also does this even need to leave the user's device?

A video editor

Say I'm building a video editor. There's a bunch of work I might want AI to help with: transcription, scene detection, classifying clips, creating embeddings, understanding what appears in a frame, searching through footage. Maybe some of that needs a giant frontier model. But all of it? Probably not.

Why upload a user's entire video to a server just to perform an operation a local model can handle perfectly well? Now you've introduced upload time, bandwidth, cloud storage, inference cost, privacy concerns, a network dependency and another piece of infrastructure I have to maintain. For what? Because API was the architecture I reached for automatically?

Local isn't automatically better

There are obvious trade-offs. Users have different hardware. Models take up space. Inference performance isn't consistent. Some tasks still need much more compute than a laptop or phone can reasonably provide. You have to think about updates, compatibility, memory and battery. All the fun stuff.

So I don't think the future is “everything runs locally.” I think the future is hybrid. Use the device when the device makes sense, and use the cloud when the cloud makes sense.

That sounds obvious when I write it down, but a lot of AI products still don't work that way. Imagine an application where lightweight tasks happen locally and instantly. Your files stay on your machine. Embeddings, basic classification and maybe transcription all happen locally. Then when you genuinely need a stronger reasoning model or heavy generation, you escalate that specific task to the cloud. Now the cloud isn't your entire AI architecture. It's one piece of it.

I like that model a lot, particularly for creative software. If I'm editing a two-hour video, I really don't want to upload the entire thing just so an AI can figure out where I stopped talking. The computer already has the video. The computer already has a GPU. Use it.

The cost of every button press

There's another reason this matters: cost. AI products can get expensive extremely quickly. Every time your user presses a button, somewhere in the world a token counter starts spinning. One user? Fine. 100 users? Still probably fine. 10,000? Okay, now we need a spreadsheet.

If a meaningful amount of inference can move onto hardware the user already owns, the economics of the product can change pretty dramatically. No per-token bill for that operation, no GPU server waiting around, no cloud processing queue. That doesn't make local inference free, because the user's machine is still doing work, but it changes who owns the compute. And I think we're going to see more products designed around that idea, especially as on-device models become more capable.

The model should follow the requirement

The part I'm most interested in isn't even “local AI” as a category. It's having more architectural choices. I don't want every AI feature I build to begin with “which API provider should I use?” I want to start with what does this feature actually need?

Maybe it's Gemini. Maybe it's OpenAI. Maybe it's an open model running on a GPU somewhere. Maybe it's a tiny model packaged directly with the application. Maybe it's a combination. The model should follow the product requirement, not the other way around.

That sounds like a small distinction. I think it's going to become a really important one.

Discussion

?

Log in to join the discussion

No comments yet. Be the first to start the discussion!

Overview