Most companies asking about AI are not starting a new product. They already have a web application or a mobile app with real users, a real database and a roadmap, and they want to know what adding AI to it involves: which feature to build, what it touches in the existing system, what it costs to run, and what happens to customer data.
This guide walks through those decisions in the order you will meet them. It is written for owners, product managers and CTOs in the US and UK, and it assumes you want a feature that survives contact with production, not a demo.
Quick answer
To add AI to an existing application, pick one narrow feature tied to a task users already do, call a hosted model API from your own backend (never directly from the browser or the mobile app), and treat the model as an unreliable external dependency: set timeouts, validate its output, cap its cost and keep a non-AI fallback. Before launch, build a test set from real examples, check what your provider does with the data you send, and release behind a feature flag to a small group first. Self-hosting an open model is worth considering only when data-control rules, very high volume or offline requirements justify the extra operations work.
Step 1: Pick a feature worth adding
The features that work best in existing products are not chat windows. They are small improvements to a task the user is already doing, where the result is easy to check and a wrong answer is cheap to correct.
Most practical AI features fall into six patterns:
| Pattern | Example in an existing product | Good fit when | Main risk |
|---|---|---|---|
| Search | Finding help articles or records by meaning, not exact keywords | Users fail to find things that exist | Returning records the user is not allowed to see |
| Summarization | Condensing a long ticket thread or call note | Staff read long text to make a quick decision | Omitting the one detail that mattered |
| Drafting | A suggested reply, product description or report section | A person reviews and edits before sending | Confident but incorrect statements |
| Classification | Routing tickets, tagging leads, flagging content | You have clear categories and past examples | Silent misrouting nobody notices |
| Extraction | Pulling fields out of invoices, emails or forms | People retype data from documents | Wrong values written into your database |
| Assistant | Answering questions over your product data | The simpler patterns are already working | Widest scope, hardest to test |
A useful filter is to ask three questions of any candidate feature:
- Can a user tell quickly whether the output is right? A draft reply is easy to judge. A summary of a 60-page contract is not.
- What happens when it is wrong? A bad tag suggestion costs a click. A wrong refund amount costs money and trust.
- Do you have the data, in a usable form, with permission to use it? A feature that needs clean historical records will stall if those records are scattered or incomplete.
Start with a feature where a person stays in control of the final action. If you are weighing an assistant or an agent that takes actions on its own, our comparison of AI agents, chatbots and workflow automation explains where each one fits.
Step 2: Call a model API or host an open model?
For most existing products, a hosted model API is the right starting point. You send a request over HTTPS, pay for what you use and have no servers to run. Self-hosting means running an open-weight model on infrastructure you control.
| Factor | Hosted model API | Self-hosted open model |
|---|---|---|
| Time to first working feature | Short: an API key and backend code | Longer: GPU capacity, serving software, monitoring |
| Cost shape | Pay per use; scales with traffic | Mostly fixed infrastructure cost, whether used or idle |
| Where data goes | To the provider, under its terms | Stays in your environment |
| Operations work | Low | High: capacity planning, upgrades, security patching |
| Model changes | Provider retires versions on its schedule | You decide when to change |
| Makes sense when | You are validating a feature or have moderate volume | Contracts or regulation forbid sending data out, volume is high and steady, or the feature must work offline |
There is a middle option. Cloud platforms offer models inside the cloud account you may already use. Microsoft states that prompts and completions for models sold by Azure are not available to OpenAI or other model providers, and AWS states that on Amazon Bedrock model providers do not have access to customer prompts and completions. For a business whose data already lives in one of those clouds, this can simplify the vendor review.
Our advice is to prove the feature with a hosted API first. If you build the integration behind your own internal interface (see the next step), moving to a different provider or a self-hosted model later is a contained change, not a rewrite.
Step 3: Where the integration sits in your architecture
The model call belongs on your server, behind your existing authentication. It should never be made directly from browser JavaScript or from the mobile app, because any API key shipped to a client can be extracted and used at your expense.
A sound layout for a web or mobile product looks like this:
- Client. The web front end or mobile app calls your own API endpoint, exactly as it does for any other feature.
- Your backend. It checks who the user is and what they may access, gathers the relevant data, and applies rate limits per user or per account.
- An internal AI service layer. One module owns the prompts, the choice of model, retries, logging and cost tracking. Nothing else in the codebase talks to the provider.
- The model provider. It receives only the data that task needs.
- Validation. The backend checks the response before it is stored, displayed or acted on.
Two points deserve attention in existing systems. First, permissions: if the feature searches or summarizes your data, the retrieval step must apply the same access rules as the rest of the application, or the AI feature becomes a way around them. Second, mobile: because the logic lives on the backend, one integration serves iOS, Android and web, whichever approach you chose in the native, Flutter or React Native decision. You can also change prompts and models without waiting for an app store release.
Step 4: Plan for latency, cost, failures and model changes
Latency
Model responses take noticeably longer than a database query, and the time varies with the length of the output. Design for it: stream text to the screen as it is generated where a user is waiting, move anything that does not need an instant answer (document extraction, bulk tagging, nightly summaries) to a background queue, and use the smallest model that passes your tests.
Cost
API pricing is based on tokens, the units of text sent and received, so cost is driven by four things: how many requests you make, how much context you send with each one, how long the responses are, and which model tier you use. Controls worth building from day one:
- Send only the context the task needs, not the whole record or conversation history.
- Set a maximum output length and per-user or per-account usage limits.
- Cache repeated instructions. Providers offer prompt caching for this; Anthropic describes it as reducing processing time and cost for prompts with consistent elements.
- Use batch processing for non-urgent work. OpenAI, for example, documents its Batch API as offering a 50% discount for jobs completed within 24 hours.
- Record usage per feature and per customer so you know what each feature costs to run.
Prices and discounts change often, so check the provider's current pricing page before you model costs.
Failures
Three kinds of failure need different handling. The provider may be slow, rate-limit you or be unavailable: use timeouts, limited retries and a fallback, which can simply be the feature hiding itself while the rest of the screen works. The output may be malformed: where you need data rather than prose, request a fixed structure (OpenAI's Structured Outputs is one example of a feature that constrains responses to a JSON schema) and still validate the values. And the output may be well-formed but wrong, which is why testing and human review matter more than any infrastructure choice.
Model changes
Hosted models are retired on the provider's schedule, and this is a recurring maintenance cost that many budgets miss. OpenAI's deprecations page describes at least six months of notice for generally available models and much shorter periods for preview models. Anthropic's model deprecations page commits to at least 60 days of notice before retirement of publicly released models. In practice this means you should reference a specific model version in configuration rather than scattering it through code, avoid building production features on preview models, and keep a test set so that a replacement model can be evaluated in days, not weeks.
Step 5: Data privacy and what you send to the provider
Every prompt is a data transfer to a third party. Before building, write down exactly which fields the feature will send, and remove what the task does not need. Names, email addresses and account numbers can often be stripped or replaced with placeholders.
Provider terms differ by product and tier, so read the terms for the exact service you are buying. At the time of writing (October 2026), the official statements include:
- OpenAI API: data sent to the API is not used to train or improve OpenAI models unless you explicitly opt in. Abuse monitoring logs are retained for up to 30 days by default, with zero data retention available to eligible customers.
- Anthropic API: by default, inputs and outputs from commercial products are not used to train models, unless you choose to share them, for example by submitting feedback. Inputs and outputs are deleted within 30 days by default, with longer retention for content flagged as violating the usage policy.
- Google Gemini API: the terms treat tiers differently. Content sent through unpaid services may be used to improve Google products and may be read by human reviewers; content sent through paid services is not used to improve products. Only paid services may be used in applications offered to users in the UK, EEA or Switzerland.
- Amazon Bedrock: retention is configurable, and some models require prompts and outputs to be retained for review by AWS as a condition of access. See the data retention documentation for the current modes.
The lesson from the Gemini example applies generally: a free tier used for a prototype can carry different data terms from the paid tier, so do not test with real customer data on a free plan.
US and UK considerations
In the UK, the UK GDPR applies when personal data is processed. The Information Commissioner's Office lists innovative technology, including AI, among the factors that can make a data protection impact assessment necessary. You will normally also need a data processing agreement with the provider and an updated privacy notice.
The US has no single equivalent law. Obligations depend on the state and the sector, so health, financial and children's data need a specific review, and existing customer contracts may restrict the use of subprocessors.
For mobile apps in both countries, Apple's App Review Guideline 5.1.2(i) requires you to clearly disclose where personal data will be shared with third parties, including third-party AI, and to obtain explicit permission first.
This is a high-level summary, not legal advice. Have counsel review your specific data flows.
Step 6: Test before and after launch
AI features cannot be tested like ordinary code, because the same input can produce different outputs and "correct" is often a matter of degree. The practical answer is an evaluation set: a collection of real, representative inputs with a description of what a good output looks like.
- Collect examples from your own data, including awkward ones: empty fields, very long inputs, other languages, angry customers, and inputs that try to instruct the model.
- Define pass criteria per pattern. Classification and extraction can be scored against known answers. Summaries and drafts need a rubric and human reviewers.
- Rerun the set whenever a prompt, model or retrieval step changes.
- Test the security cases. The OWASP Top 10 for LLM Applications is a good checklist; prompt injection, sensitive information disclosure and improper output handling are the most relevant for a feature added to an existing product.
After launch, keep measuring: how often users accept, edit or discard the output, plus error rates, response times and cost per request.
Step 7: Roll out in phases
- Prototype offline. Run the feature against historical data with nothing shown to users. This tells you whether quality is good enough to continue.
- Internal use. Put it in front of your own staff behind a feature flag. Collect corrections.
- Limited beta. Enable it for a small share of customers who have opted in, clearly labelled as AI-generated, with an easy way to give feedback.
- General release. Widen access once quality, cost per request and support volume are within the limits you set in advance. Keep the flag so the feature can be switched off without a deployment.
- Extend carefully. Only after the assisted version is trusted should you consider letting it act without review, and then only for low-risk actions.
Illustrative scenario (not a client case study): a field service company has a web dashboard and a technician mobile app. Dispatchers spend time reading long job notes. The team adds a "summarize job history" button in the dashboard, served by one new backend endpoint. Customer phone numbers are removed before the request is sent. The summary is displayed but not saved over the original notes. It ships to five dispatchers first, and the mobile app gets the same summary a release later through the same endpoint.
For technical readers
- Wrap the provider behind an interface in your own code, with the model identifier, prompts and parameters in versioned configuration.
- Treat model output as untrusted input: schema-validate it, escape it before rendering, and never pass it straight into SQL, shell commands or outbound requests.
- Enforce authorization in the retrieval query itself, not in the prompt. Instructions to the model are not an access control.
- Run long tasks on your existing queue workers with idempotent jobs, and store the model version alongside any generated content.
- Log request metadata, token counts, latency and outcome. Decide deliberately whether to log prompt contents, since those logs then hold personal data.
- Set spend alerts and hard limits at the provider and per-tenant quotas in your own application.
Frequently asked questions
Do we need to rebuild our application to add AI?
Usually not. Most AI features are added as new backend endpoints and a small amount of interface work. Rebuilding is only on the table when the existing system has no usable API layer or the data the feature needs is inaccessible.
How much does it cost to add AI to an existing app?
There are two costs. The build cost depends on the feature's scope, the state of your data, how many systems it touches and how much testing it needs. The running cost depends on request volume, how much text is sent and returned, and the model tier. A narrow feature such as classification is far cheaper on both counts than an assistant.
Will our customer data be used to train the AI model?
It depends on the provider and tier. The major commercial APIs state that business API data is not used for training by default, but free tiers and consumer products can have different terms. Check the current terms for the exact service you use, and confirm retention periods as well as training use.
Should we fine-tune a model on our own data?
Rarely as a first step. Good instructions plus retrieval of the right records at request time solves most product use cases and is much easier to update. Consider fine-tuning later if you have a stable, high-volume task and a large set of quality examples.
Can the AI run on the phone instead of a server?
Small on-device models exist and suit some narrow tasks, but capability varies by device and operating system version. For most business features a server-side integration is more consistent across iOS, Android and web, and easier to update.
How long does it take to add an AI feature?
A working prototype is quick. The time goes into what surrounds it: permissions, data preparation, the evaluation set, failure handling, privacy review and a staged rollout. Plan the timeline around those items, not around the model call.
Conclusion
Adding AI to an existing product is mostly ordinary engineering around one unusual dependency. Choose a narrow feature with a checkable output, keep the model call on your backend behind a single service layer, send the minimum data, test against real examples, and release in stages with a switch to turn it off. Teams that do this get a feature they can maintain when prices, models and regulations change.
A sensible next step is a one-page feature brief: the task, the data fields involved, who reviews the output, and what "good enough" means. If you are bringing in an outside team, our guide to outsourcing software development from the US or UK covers how to assess one. Entrant Technologies builds web applications, mobile apps and custom software, and if you would like a second opinion on a feature brief you can get in touch here.