Shipping on-device AI with Apple's Foundation Models

2 min read by Michal Ferák

journeybot builds packing lists for a specific trip: destination, dates, weather and planned activities. Since iOS 26, Apple’s Foundation Models framework lets apps call the same on-device language model that powers Apple Intelligence. We use it so that trip plans never leave the user’s device. Here is what we learned getting it to production.

Treat availability as a feature, not an error

The on-device model isn’t available everywhere. The device may not support Apple Intelligence, the user may have turned it off, or the model may still be downloading. SystemLanguageModel.default.availability tells you which, and each case deserves a different response.

We designed journeybot so the app is complete without the model. Packing lists are first assembled from rules based on weather, trip length and activities. When the model is available, it refines and personalises that list. When it isn’t, the user still gets a good list, and nothing on screen looks broken.

Ask for structure, not prose

Free-form text is hard to put into a checklist UI. The framework’s guided generation lets you describe the output as a Swift type with the @Generable macro, and add @Guide descriptions to its properties. The model then produces an instance of that type directly.

This removed a whole class of bugs. There’s no JSON parsing, no regex over model output and no “the model added a friendly sentence before the list” surprises. The type is the contract between the model and the UI.

Keep prompts small and specific

The on-device model is much smaller than hosted frontier models and has a limited context window. It’s good at focused tasks and gets worse as you pile on instructions. What worked for us:

  • One task per session. Packing suggestions and weather summaries are separate calls with separate instructions.
  • Give it facts, not raw data. We summarise the forecast into a few lines before passing it in, rather than handing over the full weather payload.
  • Short instructions in plain language, with an example of the tone we want.

Stream for perceived speed

Generation on-device is fast but not instant. Streaming partial results with streamResponse lets the list fill in item by item, which feels quicker than a spinner followed by a full list, even when the total time is the same.

Test like it’s a dependency you don’t control

Model behaviour can shift between OS releases. We keep a set of sample trips — a beach week in July, a business trip to Oslo in January, a hiking weekend — and run them after every OS beta to check that suggestions still make sense. It’s a lightweight evaluation set, but it has caught regressions before users did.

Why on-device was worth it

Running the model locally means no API costs that scale with users, no server to maintain and a privacy story we can state in one sentence: your plans stay on your device. For an app built by a small team, that combination is what made AI features viable at all.

If you’re considering on-device or hosted AI for your product, we help teams make that call and build it.

Keep reading

  • How an AI-native software house actually works

    What changes when AI agents write much of the code, what doesn't, and why senior judgment matters more than ever. A look at how Loam delivers software.

  • Accessibility audits that end in fixes

    How to run a WCAG accessibility audit on an e-commerce site so that it leads to shipped fixes, not a spreadsheet nobody opens. With the European Accessibility Act in mind.