The data boundary

Does running the model yourself actually make it safer?

Moving the model in-house changes who owns the infrastructure. That is not the same as lowering the risk.

Short answer

Where a model runs settles two questions: where data goes, and who is responsible for the infrastructure. It settles none of the questions that decide how large the damage is. Prompt injection, over-broad permissions, unsafe tools and architectural mistakes are exactly the same on both sides. The choice should be made on classes of data, on requirements you are obliged to meet, and on whether your team can operate a runtime, rather than on the slogan that in-house means safe.

#A decision made in ten minutes

The conversation takes ten minutes and sounds convincing. Legal is worried about personal data, the technical lead suggests running the model in-house, everybody agrees that this is safer. A line appears in the notes: we are moving to our own model.

A month later the company has a local model. It also has an administrative token for the internal API, unrestricted outbound internet, write access to the database and a tool that sends email. The personal data genuinely does not reach an external provider. Everything else that could have happened still can.

I am not arguing with the decision. Running locally is sometimes the right choice and often the only possible one. I am arguing with the conclusion that arrives with it: that the security question is now closed. A substitution happened there. The question of the data boundary was answered, and the box beside the question of damage was ticked.

#Two different questions

Any system with a model in it has to answer two independent questions, and where the model runs answers only the first.

  1. Where data ends up and who owns the infrastructure

    Which data the model sees, where it is processed, who else can reach it, what is written to logs, how long that is kept, under which contract and in which jurisdiction.

    This is the question the cloud-or-own-server choice really does affect.

  2. What the agent can do

    Which tools are connected, what rights they hold, whether volume is capped, what happens without confirmation, what cannot be undone, how many minutes before a person hears about a problem.

    Where the model runs has no effect on this at all. The answer lives entirely in the architecture of your application.

The second question decides the size of the damage. The first decides who you trust with data. They get confused because both are called security, and because the first is settled by a purchase or a move while the second is only settled by work.

#What genuinely changes

None of this is an argument against running locally. It has real advantages, and they are worth naming precisely so that nothing extra gets attributed to them.

Five risks, two places to run

The top two rows change. The bottom three, the ones that decide how large the damage is, are identical in both columns.

  • The data boundary. Running locally means the contents of your requests do not leave your infrastructure. That is a real and checkable property, and for some classes of data it decides everything.
  • Contractual and legal requirements. Some obligations cannot be met by a sentence in a provider privacy policy. If your contract or your regulator requires processing inside a defined perimeter, the question is closed before the discussion starts.
  • Predictable behaviour. A cloud model changes when the provider decides, and the answer to the same request can change with it. Your own model changes when you change it. The price is that improvements also arrive only when you bring them.
  • Who operates it. This is the advantage that is easy to mistake for a drawback, and the other way round. In the cloud, patching, isolation and availability are not your work. In-house they are entirely your work: updates, monitoring, access, on-call, network segmentation.

That last point is worth costing in advance. Running locally hands your team a new thing to operate. If there is nobody to patch it and nobody to look at a graph at night, you have traded a known supplier risk for an unknown one of your own.

#What stays exactly the same

Here is the part that matters. None of the following changes when the model moves, and this is the list that decides what a mistake costs.

There is a one-move way to check it. Imagine the model is already inside your network and nothing else moved: the same tools, the same tokens, the same caps. Now walk the list of things you are afraid of and mark the ones that just became impossible. Usually exactly one entry goes to zero, the leak of request contents to an outside provider, and the rest stay in place word for word.

Risks on both sides of the boundary
Risk Changed by where it runs What actually bounds it
An instruction arriving inside data No. A local model reads text as meaning in exactly the same way. Separating data from instructions, tool permissions, confirmation before the irreversible.
Over-broad permissions No. An administrative token stays an administrative token. Rights over one object, an expiry on permission, separate credentials per agent.
Dangerous tools No. Being able to send an email or delete a row has nothing to do with where inference happens. The list of tools, a cap on volume, a separate path for bulk operations.
A leak through the output Partly. The provider will not see it; the recipient of the email will. Rules about what the agent may send and to whom, and a check before sending.
Logs and retention Yes, but not by itself. Your own logs also hold data and are also read by somebody. A retention policy, access to logs, removing identifiers.
A mistake inside a chain of steps No. A cascade is built out of architecture, not out of a model. A depth limit, idempotency, a re-check before the irreversible step.

#How to choose

Four questions, in this order. The first two usually settle it, and the fourth is rarely reached.

Scene Invented example — not a customer

A company connects a model to its customer conversations. The whole customer profile goes into every request, because that was easier to write: one object, one line of code. Of that profile the model needs two fields, the plan and the renewal date. The rest, including the address, the phone number and the payment history, travels with them every time.

While the meeting debates moving the model in-house, the most sensitive part of the problem is solved by one change on the application side. This is the usual order of events: the question of where to run arrives before the question of what is being sent at all.

  1. Describe classes of data, not data

    Split what ends up in requests into groups: public, internal, personal data, payment details, medical or otherwise special. The decision is then made by the most sensitive group the model genuinely needs.

    It often turns out the model needs one group and is being sent another. That is the cheapest fix in the whole project.

  2. Write down the requirements you are obliged to meet

    A contract with an enterprise customer, sector regulation, a promise in your own privacy policy. A requirement either exists or does not, and it is not discussed in terms of convenience.

    If one exists, you can stop reading: the place has been chosen for you.

  3. Assess your ability to operate it

    Who updates it, who watches the load, who answers at night, who manages access, what happens when one machine fails. The honest answer of nobody is more common than it sounds.

    Running locally without operating it is worse than a cloud with a mediocre contract, because at least somebody maintains the second one.

  4. Design the agent limits separately

    Permissions, caps, confirmations, rollback, logs. This work is identical in both cases, and it is the work that decides the damage. It cannot be left for later, because later almost never has time in it.

    Do only this step and change nothing about where the model runs, and the damage goes down. Move the model and do nothing else, and it does not.

#Our own answer

1ADK is built so that this choice stays yours. The analysis is run by an agent on your side, the one you already use, and what arrives in the service is structured findings and evidence. 1ADK does not require access to your repository and does not require uploading source code to it. The architecture is not tied to one model provider.

What we are not saying

We do not claim that with this arrangement nobody sees the code at all. A cloud coding agent reads your code under its own provider terms, and that is a separate decision, which is yours. Our boundary says something narrower: what 1ADK itself cannot do. The difference between those two statements is exactly what is worth demanding from any supplier.

#What gets confused

The data does not leave our server, so we are safe.

Instead That is a statement about confidentiality, not about safety. It says nothing about what the agent can do to your database, your money and your mailbox.

A local model is cheaper because there is no token bill.

Instead Cost it fully: hardware, updates, monitoring, on-call, downtime and engineer time. Sometimes local is cheaper, but that is the result of a calculation rather than a property of the approach.

The model is ours, so injection is not our problem.

Instead Injection is a property of how a model treats text, not of where the server stands. The defence is separating data from instructions and bounding tool permissions.

One supplier decision will handle this.

Instead A supplier decision closes the data-boundary question. The damage question is closed by separate work on permissions, caps and confirmations, and the size of that work does not depend on the supplier.

#What this article does not claim

What this does not do

  • It does not claim the cloud is safer. For some data and some requirements it is simply not applicable.
  • It does not claim running locally is safer. It moves responsibility rather than removing it.
  • It does not assess model quality and does not recommend a particular provider.
  • It does not replace legal advice: which data may go where is settled by contract and regulation, not by architecture.
  • It does not cover training on your data. That is a separate question, answered by provider terms and settings rather than by where inference happens.

Take your most visible piece of AI automation and answer one question: if the model moved in-house tomorrow, which of the risks in the table above would change? If the answer is none of the bottom three, then the work that actually reduces damage has not started yet.

First find out where your data already goes.

The decision about where to run a model rests on which data travels where in your system today. 1ADK shows that: data stores, integrations, interfaces, and the evidence under every finding.

Build your project map Our boundary and what we do not promise