A phone where you’d barely need to open your applications at all. A strategy years in the making, anchored by a rebuilt Siri and a new generation of chips.
Find that lost reservation buried in your emails, pull up photos from a trip, edit them, share them. That’s what the new assistant promises. Put like that, you might think: okay, cool, a few practical shortcuts, save me three taps, sell another phone. But extend the logic.
You’re Still the Delivery Driver of Your Own Phone
Right now, to get a result, you connect the dots between your own apps yourself. You search for information here, copy it, open something else, start over. You are, in a proper sense, the delivery service of your own phone. If an AI can do that job for you simply because you said what you wanted, or better yet, because it understood automatically, then a whole chunk of that navigation becomes unnecessary. And the company that orchestrates all of it could capture far more than a few seconds of your attention and a healthy pile of dollars.
That’s the hypothesis, at least. By controlling how AI functions on your behalf, Apple could capture a major share of its value. Obviously, between that destination and a Siri that actually delivers, there’s still a road to travel.
Because to hand your daily life over to a machine, you need to count on it all day long, fast enough, without killing your battery, and at a sustainable cost. So all this intelligence at your service, where does it actually run?
Where Does Siri AI Actually Run?
As you probably know, a model’s life doesn’t end when training is over. Training builds capabilities. Inference is when you use them: understanding your request, analyzing an image, producing a response. And every single use consumes compute and energy. A single request can even trigger multiple model calls as the assistant searches, compares, and cross-checks information. Yes, improving models is still essential. But as AI enters daily routines, the importance shifts toward the cost and conditions of its repeated use.
To give you a sense of scale, Deloitte estimates that two-thirds of all AI compute power in 2026 will be dedicated to inference, not training. And even though technical improvements make each operation cheaper, their sheer multiplication keeps pushing the bill higher.
For you as a user, there’s also a much more immediate constraint: time. If the assistant makes you wait longer than the manual steps it was supposed to spare you, that makes zero sense. If you need to rephrase your request three times, it doesn’t work either. Useful performance here means completing the task with minimal friction.
Going fast enough, pulling the right information, using the right resources. And that quality of service has to hold a massive scale. Over a billion iPhones are active worldwide, roughly one in four smartphones on Earth. Imagine those operations repeated from morning to night, even on a fraction of those devices.
On top of that, not all requests need the same level of resources. Finding a booking number in an email is one thing. Analyzing 300 pages of contracts, invoices, and correspondence to spot contradictions is another. You need to process more information, cross-reference it, and multiply the checks.
Always deploying the heaviest resources would be pointlessly expensive. But limiting every request to the same small resources would cripple the assistant’s capabilities.
That’s why where the computation happens becomes a strategic question. You can bring the power to your data, or send the necessary information to where more power lives.
In the first case, the model rushes to your phone. It can handle a request with no round trip to a server, keeping data on the device. For something like dictation, this matters a lot. You want words appearing as you speak. But your phone has limited memory, a battery, and heat to dissipate. The more you ask it to compute, the more those constraints bite. And the assistant still needs to leave you enough battery life to get through the day.
On the cloud, on the other hand, you can throw much more memory and power at the problem and run larger models. Fantastic. But in return, you need a connection; you have to route information back and forth, wait for the response, and someone has to pay for those servers. And let’s not forget: if we’re touching your emails or family photos, the question of who can access that data matters quite a bit.
So each approach has its strengths. Why send a simple request to a server if the phone can handle it correctly, on the spot, immediately? And why limit a complex task to what fits in your pocket?
Apple understood this. It combines both.
Apple’s Hybrid Bet: The A20 Pro and Private Cloud Compute
Since 2024, Apple Intelligence has been pairing local models with cloud reinforcement through what they call Private Cloud Compute. The system evaluates each request’s demands and can call on a remote model when more power is needed. From your end, it’s completely transparent. You don’t even know what’s happening behind the scenes.
To protect the data that gets sent, Apple designed this cloud with specific guarantees: information must only be used to process the request, then deleted, and never accessible to Apple employees.
And then there’s the chip angle. Because Apple controls the entire stack- hardware, operating system, and models-it can optimize them as a unit. On the A20 Pro inside the new iPhones, the company is announcing 50% more memory bandwidth compared to the A19 Pro, among other upgrades. In plain terms, more data can flow to feed the calculations. That’s what this generation brings to Apple’s bet: more muscle to run AI on the device itself.

And from a business perspective, because let’s be honest, this is still a business, every task Apple processes locally through your iPhone is potentially one fewer call to its servers. The company is then leaning on compute power already bought by its customers. For the heaviest workloads, its architecture extends to Nvidia GPUs running in Google Cloud. So you end up with two types of chips, Apple’s and Nvidia’s, working together to serve different needs within the same device.
Why Apple Paid $1 Billion a Year for Google’s Brain
Apple now has much of the machinery to run this assistant at scale. But you still need models capable of understanding requests and executing them correctly. And for that, Apple turned to an old friend: Google.
On January 12, the two companies announced a multi-year deal. Gemini’s technologies would serve as the foundation for Apple’s next generation of models, particularly to power Siri. For a company that has historically made controlling its own products a core strength, that’s nothing. Apple is accepting that it needs to go outside for part of the intelligence it requires.
And remember, Apple had already committed on this front. Back in 2024, it presented its own language models optimized for everyday tasks on its devices and in the cloud. But here’s the thing: having the customers doesn’t automatically mean you can build the best model. You also need to master the full engineering stack required to produce a reliable system.
In June 2025, Craig Federighi explained that the first architecture for the new Siri didn’t meet the quality bar they expected. The company had to rethink its approach and push back the features it had promised.
If you ask me, I think Apple is trying to push forward despite its struggles. Lean on Google to strengthen the intelligence of its products while leveraging an advantage it already owns: the relationship with users. Because a model, however powerful, still needs to find its place in your life. Discovering a service, deciding to try it, understanding what it can do, then building the habit of coming back to it. Apple already has a daily point of contact: the iPhone. And around it, the Mac, iPad, and the rest of its ecosystem. Integrating the assistant directly into its products gives it a shot at becoming useful exactly when you need it, while you’re reading a message, looking at a photo, or preparing a document.
And that access to users already has a price tag. In 2022, Google paid Apple $20 billion as part of their search agreement, notably to remain Safari’s default search engine. On the flip side, for Gemini, reports put the payment at roughly $1 billion per year from Apple to Google. Google pays to access Apple’s users. Apple pays to access Google’s model.
The Real Advantage Isn’t the Model
But Apple’s advantage goes far beyond distribution. Take a request like “send Guillaume the file from this morning’s meeting.” To execute that correctly, the system needs to identify the meeting, find the right file, figure out which Guillaume you mean, and have an authorized way to send the document.
Language comprehension is only part of the job. You need to connect it to context and to the actions available across your applications. And that’s exactly what Apple is building. Through the App Intents framework, developers can make information and functions from their apps accessible to Siri. The assistant can then search for content and trigger an action the app has made available.
The model provider can stay essential, but the company that receives your request, retrieves your context, and organizes its execution occupies a different strategic position entirely. My reading of this is that Apple is trying to control that passage point. The moment where intelligence becomes a useful action for you. It’s this position that could let it turn a technical advance into an everyday reflex. You have something to do; you go through its assistant.
And if that reflex takes hold, it could change far more than how we ask questions.
What Happens When Apps Become Invisible
Why go through every app yourself to find a result when you can ask for it?
Okay, I’m entering speculative territory here, and I’m perfectly happy to be wrong about this. But I find this trajectory coherent, and I could absolutely see it playing out.
Imagine you want to organize a dinner with three families of friends. You need to find availability in the group chat, search for a restaurant, make sure it works for everyone, book it, then send the address and time back to the group. You’ve been there, I’ve been there, and it’s not the fun part.
Today, you coordinate this minor operation across several apps. Following Apple’s example, you could say: “Look for a restaurant in this neighbourhood on Friday night that serves vegetarian options and is under €40 per person.” The assistant would match your request against the information you’ve given it access to, compare possibilities, present you with a choice, and after your go-ahead, handle the reservation, the route, and the message. And if you add “actually, make it 8:30 instead,” it would know exactly what you’re referring to.
One conversation would be enough to pilot an operation that used to require multiple searches and manipulations. Voice would become a shared interface across your services. The screen would review proposals, verify a detail, confirm a purchase. A huge portion of menus and back-and-forth could disappear.
That’s what I mean by a world where interfaces become almost invisible. When you grab a glass, you focus on one simple gesture without consciously detailing every movement of your fingers. You could imagine information technology delivering that same simplicity. You state the desired result and the system coordinates the necessary operations. All the complexity still exists, but it takes up less of your attention. It almost disappears from conscious awareness.
Apps would keep providing their services in the background. You’d use their functions without entering each one.
And actually, this separation between an app and its functions is already at the heart of Apple’s developer tools. Since 2024, the company has been showing developers how to make their actions and content accessible elsewhere in the system, notably through Siri. The next step would be connecting these capabilities intelligently enough that you could delegate an entire goal. And at the end of that road, the logic could spare you even the request itself.
Among the functions announced this year, Apple describes a system that can automatically surface your booking number when you call an airline. It identifies the company you’re calling and retrieves the information from your emails. You don’t even need to ask it to search.
In our dinner scenario, the assistant might notice that the agreed time in the group messages overlaps with an appointment in your calendar, then suggest an alternative before you’ve even spotted the problem. Every time context suffices to understand a need, one more step can vanish.
Who Chooses the Restaurant?
Apple has multiple touchpoints to make this assistance accessible. iPhone, Apple Watch, AirPods, CarPlay. The new Siri is announced across all of them.
But now let’s imagine this works, that it’s widespread, that it’s your daily reality.
Who picks the restaurant? You gave a budget, a neighborhood, a few preferences. Several possibilities still fit. Why does the assistant show you three instead of ten? Why does it go through this booking service? What criteria determine what lands in front of you?
You can keep the last word while letting the assistant choose which options you get to pick from. And this question goes way beyond my dinner-with-friends example. It’ll be the car that gets sent to you, the product you’re nudged to buy, the information selected to answer your question. Whoever organizes those choices could gain considerable influence over the companies that get your money and your attention.
For a service, the commercial stakes could shift to the assistant proposing it the moment a need appears. For Apple, controlling this meeting point between an intention and its execution could become an unprecedented source of value.
And the most interesting part is how this change might install itself. I don’t see a big overnight reveal, some dramatic break between Sunday and Monday morning. At first, you’ll try it to see. Then you’ll come back because it was genuinely useful. Until the day when digging through menus yourself feels completely pointless and unnecessarily complicated. Adoption could happen exactly like that, through a succession of small services rendered.
In this scenario, you’d be using AI all day long. But the model behind it? You wouldn’t care at all. And I can tell you that where it runs its calculations will be the absolute last of your concerns. Just like the chips in your computer: you can benefit from what they make possible without knowing their architecture. What matters is what you do with them.
AI will end up on the same path. Becoming an ordinary capability of the surrounding objects. Something we eventually expect from them.
The Lock-In Nobody’s Talking About
That’s how the fresh pieces of Apple’s strategy make sense together. Local compute and cloud need to make assistance available under acceptable conditions. Models need to understand the request. System integration needs to let them retrieve context and act. Apple is presenting the new Siri as the convergence of all these capabilities.
For the user, all this complexity would boil down to one experience: I asked, and it’s done. If Apple can repeat that experience often enough, it’ll also reinforce a relationship that goes beyond the phone itself.
When it’s time to choose your next device, it won’t just be about wanting to stay in Apple’s ecosystem. It could also, and especially, be about wanting to find that familiar assistance again. Add a recognizable voice, a personality, the memory of your habits. Some people might perceive this assistant as a companion.
Almost a friend.
Changing ecosystems might feel like abandoning a familiar presence that recognises you.
For Apple, brand loyalty would then take on an emotional dimension. The clearest sign of success would be a shift in reflex: face a task, start by asking the assistant to handle it.
But Let’s Not Get Ahead of Ourselves
Coming back to the new iPhones, they mark a milestone in this strategy. But reality still has a few things to say. In iOSWorld, a benchmark published in June, the best configuration evaluated only achieves 37% success on tasks involving multiple applications. The new Siri wasn’t among the systems tested, but it illustrates how hard multi-app coordination actually is.
For this to work, delegating needs to cost less effort than doing everything yourself. Verification included.
If Apple’s bet pays off, the smartphone could paradoxically take up less and less space in our gestures while taking up more and more space in our decisions. And if that happens, then the next great iPhone progress might just be measured by everything you no longer need to do on it.
Thanks for reading. Let me know your thoughts in the comments. Don’t forget to follow us and subscribe for more analysis. If you like space and physics innovations, follow The Nov Science for news related to science.







