Start with a responsibility, not a personality

An AI agent is a system that uses a model to select or execute steps toward a defined task. The important design question is not what to name it. It is what responsibility it can safely accept. “Prepare an intake record for review” is a clearer scope than “run customer service.”

Write down the trigger, required inputs, allowed tools, intended output and conditions for stopping. Include examples where the right response is a clarification, an escalation or no action. This gives both engineers and business owners something concrete to evaluate.

Give tools narrow contracts

A tool should represent a specific business action with a typed input and a predictable response. The model does not need unrestricted access to your database or arbitrary code execution to look up an appointment. It needs a scoped operation that checks authorization independently.

Validate tool arguments on the server. Confirm that the requested record belongs to the right account and that the action is permitted in its current state. Treat any instructions found inside retrieved documents, emails or web pages as untrusted data.

Separate knowledge, memory and records

Knowledge retrieval supplies relevant documents for the current question. Conversation memory preserves recent context. A system of record stores authoritative business state. Mixing them can cause stale information to be treated as current fact.

Define what is remembered, for how long and on whose behalf. A corrected address belongs in an approved customer record, not only in a model’s conversation history. A deleted or restricted document must stop appearing in retrieval.

Make approval meaningful

Before a consequential action, show the person exactly what will happen: the recipient, changed fields, document or message, and the supporting context. Approval should refer to this specific proposal. If the proposal changes, the old approval should not silently apply.

A handoff must also work when no reviewer is available. Decide whether the task waits, expires or returns a clear status to the requester. Log both the decision and the executed action without collecting unnecessary sensitive information.

Evaluate complete tasks

Test more than whether an answer sounds plausible. Did the agent choose the correct tool? Did it retrieve an authorized source? Did it stop when evidence was missing? Did a retry create a duplicate action? Include failed dependencies and adversarial input in the evaluation.

Start with a representative test set and a simple baseline. Measure task completion, unsupported claims, escalation, latency and usage cost separately. A model change should pass this evaluation before it reaches the live workflow.

A sensible first release

Choose one bounded workflow, limited tools and a small group of users. Keep high-impact changes behind review and observe the exceptions. The early goal is to understand real behavior, not to claim complete autonomy.

Bring a handful of real, appropriately redacted examples to discovery: successful cases, ambiguous requests and cases your team would decline. They reveal more about the system you need than a long list of AI features.

Turn the thinking into a project

Use the project planner to describe your objective, integrations and constraints. The result is a discussion brief, with no invented price or delivery commitment.

Plan your project
Explore the related engineering capability