Running AI agents on your own infrastructure
An agent that works against a sample database has answered one question: can it complete the task? Running it against company data raises several more. Which systems can it reach? Whose permissions does it use? What happens when a source is unavailable, or when the same event arrives twice?
These decisions shape the system around the agent. Here is how we would work through them before putting a workflow into production.
Decide where model requests go
Running the agent platform on your infrastructure does not, by itself, keep model inference there. The platform and the model are separate deployment decisions.
With Speculos, the platform runs in your infrastructure; on AWS, that means an AMI in your account. You can host models inside your environment or connect an external model service through your own provider account and keys. In the second arrangement, the content included in a model request goes to that provider.
Before deployment, identify the hosting account, the network paths to each data source and the model endpoint. Record which data may be included in model requests. If inference must stay inside your network, model hosting and its compute requirements belong in the deployment plan too.
There is no useful universal instance size for this. A scheduled document review using an external model has different requirements from a locally hosted model serving concurrent investigations. Size the system against a representative workload, including the slowest tools and expected concurrency.
Connect private systems without sharing credentials
An integration needs more than an endpoint. It needs a defined set of operations, an identity to perform them and an access policy that still applies when an agent is shared.
For each source, answer four questions:
- What can the agent read, and can it write anything?
- Whose identity authorizes each tool call, including scheduled runs?
- How are results restricted to the right customer, team or account?
- What happens when access is revoked or credentials rotate?
In Speculos, administrators connect sources at the company level and grant access. People build against the sources they have been granted without pasting credentials into a prompt. Access is checked when the deployed agent is used. Our data-governance article explains why those checks matter after someone shares an agent.
For an internal API, start with a narrow tool interface: explicit inputs, bounded results and a clear error response. A customer identifier supplied by an agent should be checked against the caller's authorization before any records are returned. Prompt instructions alone cannot enforce that boundary.
Speculos MCP is a separate, hosted offering for controlled database access. It exposes PostgreSQL through MCP, with column rules and query logging. Choose that arrangement deliberately: it runs on the Speculos side, unlike the platform installed in your infrastructure.
Give each agent a bounded job
Some workflows need more than one agent. A research workflow might separate finding documents, extracting facts and checking a conclusion. Those divisions are useful when each step has a clear input, output and reason to exist.
Define the handoff before adding another specialist. Include the evidence it should return, the conditions under which it should stop and who handles an incomplete result. Keep customer records separate from reusable instructions so a specialist does not carry one customer's context into another run.
A coordinator also needs limits: timeouts, retry counts, a maximum number of delegated tasks and a budget. If the workflow can create a specialist dynamically, validate its configuration and tool access before activating it. These are design requirements for that workflow, not permissions an agent should grant itself.
Keep a trace that explains the outcome
A final response is not enough to investigate a failure. A useful run record connects the triggering event to the agent versions, tool calls, evidence and final decision.
Capture the inputs and outputs needed to reproduce a decision, while applying access and retention rules to the trace itself. Logs can contain the same sensitive information as the underlying sources. Record tool errors, retries, latency and model usage so an engineer can distinguish a bad conclusion from an unavailable dependency.
The record should also explain why nothing was sent. Suppressing a duplicate, finding no applicable evidence and failing to finish are different outcomes.
For example, a vendor-risk workflow could read a vulnerability notice, check authorized customer and vendor records, and assess whether the affected product and version are present. Its outcomes might be a supported notification, suppression of a previously reported finding, or review when evidence is incomplete. That is an illustrative workflow; the same approach applies to research and document review.
Test the workflow before expanding it
Start with one workflow whose output can be assessed by the people who use it. Keep examples of correct results, missing evidence, denied access and duplicate events. Include a revoked permission and a failed tool call, not just successful runs.
Measure completed work and review burden alongside cost and runtime. A system that produces more findings but creates more checking for the team may not be an improvement.
Speculos brings governed data access, agent building and execution onto company infrastructure. Our engineers can help with the initial implementation alongside your team. Book a demo to work through the deployment and data requirements for your first workflow.
