Apple’s Foundation Models framework now gives developers a single Swift API for Apple Foundation Models on device, Apple Foundation Models on Private Cloud Compute, and other language-model providers that conform to Apple’s protocol. In WWDC26 developer sessions and Apple Developer materials, Apple says supported local inference carries no per-inference cost, while many App Store Small Business Program developers can access an Apple server model on Private Cloud Compute at no cloud API cost, subject to download eligibility and per-user daily limits.
Apple has long sold on-device processing as privacy. The newer developer message puts cost beside privacy. A feature that summarizes a journal entry, labels a receipt, suggests task tags, extracts details from a note, or rewrites a short passage may run dozens of times per person. If every tap calls a rented cloud model, each successful interaction also adds variable cost.
If the same job runs on hardware already in the user’s pocket, or through Apple-managed capacity with usage limits, the app maker can think about the feature differently. A capability that would be too expensive to offer without caps may become a normal part of the interface.
Strictly speaking, most app developers are not buying and owning AI compute in Apple’s model. The compute is already in the customer’s device, or it is delivered through Apple’s Private Cloud Compute and iCloud system. But for product economics, the effect can be similar: the marginal cost of a supported local request can drop close to zero for the developer, while a cloud API request remains metered.
That contrast is visible in public pricing pages from OpenAI and Anthropic, where API costs are still organized around input tokens, output tokens, cached tokens, tool usage, batch processing, or other usage units. That billing model is clear and flexible, but it ties engagement directly to cost of goods sold.
Token-free is not cost-free
Apple’s own materials draw the boundaries. Apple’s on-device Foundation Models now include multiple variants. AFM 3 Core is a roughly 3-billion-parameter model, while more capable hardware can run AFM 3 Core Advanced, a larger sparse model. Apple positions these on-device models for tasks such as summarization, extraction, classification, text refinement, short dialog, structured output, and tool calling rather than as replacements for server-scale models used for broader world knowledge or advanced reasoning.
Availability is another cost. Apple’s September support page lists Apple Intelligence support across newer hardware, including iPhone 16 models or later, iPhone 15 Pro and iPhone 15 Pro Max, iPhone Air, iPad mini with A17 Pro, MacBook Neo with A18 Pro, iPad and Mac models with M1 or later, Apple Vision Pro, and supported Apple Watch models when paired with an Apple Intelligence-enabled iPhone. Apple also lists language, region, operating system, and storage requirements.
Server-side Apple Intelligence features have their own constraints. Apple says certain features that rely on server-side models are subject to daily usage limits, with limits varying by feature, request complexity, system demand, and other factors. Apple also says expanded access to those features will be available for a fee in the future.
Those limits change the break-even calculation. A local AI feature may still need a fallback for unsupported hardware, unsupported regions, long documents, fresh web knowledge, stronger reasoning, or business rules that require centralized logging and review. The fallback is often a cloud model, which brings token pricing back into the cost model.
When local compute starts to win
The strongest case for local AI is a frequent, bounded task on private or personal data. Summaries, label suggestions, extraction, proofreading, tone rewriting, receipt categorization, note organization, and lightweight app commands fit that pattern because the model receives short context, the answer can be checked by the app or the person, and the task may run many times.
TechCrunch’s October 2025 roundup of early Foundation Models apps fits that pattern. It found developers using Apple’s local models for story generation in a learning app, spending insights and categorization, word examples, task tagging, journal highlights, recipe steps, contract summaries, workout summaries, and other app-level features. The reporting also framed many of those features as quality-of-life upgrades rather than major workflow changes, which lines up with Apple’s own guidance about the size and intended use of the on-device model.
A simple break-even test is repeated prompt volume multiplied by API price, plus monitoring and retries, compared with added engineering, QA, support, battery constraints, and the share of users whose devices actually qualify. If the first side grows with every use and the second side is mostly fixed, local inference starts to look cheaper as usage rises.
When rented AI still makes sense
Cloud APIs still make sense when model quality or infrastructure matters more than per-request cost. That includes advanced reasoning, large-context document review, live web knowledge, multi-platform products, heavy tool use, multimodal generation, strict enterprise operations, or workflows where a better answer saves enough money to justify the meter.
Cloud billing can also be the simpler architecture for teams that do not want to tie a feature to Apple Intelligence availability. A business with Android users, Windows users, web users, older iPhones, or regulated workflows may decide that a single cloud model with centralized policy, logging, and support is worth paying for.
That decision is not only technical. It affects pricing, margins, support, privacy promises, and product positioning. A small app with generous AI features may favor local inference to avoid usage caps or paid tiers. A business system where accuracy, auditability, and cross-platform access matter may accept token billing because the feature is valuable enough to support it.
Apple’s platform bet is cost control through integration
Apple is not eliminating the economics of AI inference. It is moving more of that economics into the platform. The device price, chip design, operating system, iCloud account, App Store program status, model limits, and usage caps all become part of the cost structure.
For Apple, that makes AI a reason to buy newer devices and stay inside its ecosystem. For developers, it may reduce the fear that a successful AI feature will punish margins with a bigger token bill. For users, it can keep more personal data on device or inside Apple’s Private Cloud Compute design, though only within the devices, regions, languages, and daily limits Apple supports.
The lesson for businesses experimenting with AI is similar to the one behind careful AI use in content workflows: lower generation cost does not remove the need for judgment, QA, and a clear reason for the feature to exist. Tech Help Canada has made that point in its guide to using AI tools for SEO without publishing low-value content, and the same standard applies to product features. Cheap AI that produces weak output is still expensive if it creates support work, trust issues, or unused features.
The better cost model is no longer “AI feature or no AI feature.” It is a routing decision: local by default where the task is bounded, Apple Private Cloud Compute where deeper capability is available and limits fit the product, and cloud APIs where capability, reach, or control matter more than the meter.

Tech Help Canada Staff researches, writes, and reviews practical content for business owners and professionals. Our coverage spans business, marketing, SEO, technology, and the tools and systems people use to grow and operate online. We focus on clear, useful information backed by research, hands-on experience, and editorial review. Learn more about our team and editorial standards. Need help with something? Contact Us







