ASSESSMENT AND IMPLEMENTATION

AI agents for business, with a specific task and explicit limits.

An AI agent can consult information and use tools to complete a task. Stolen Orbit designs these workflows around bounded objectives, limited access and approval points, so your team can understand, check and stop what happens.

Let’s discuss your process

What is the difference between a chatbot, workflow and AI agent?

This distinction helps design the work. More autonomy is not a success criterion in itself. If a deterministic workflow covers the need, there is no reason to add an agent.

A practical distinction for choosing a project
ApproachWhen to assess it
Chatbot or assistantYou need a conversation or proposed text; a person decides and carries out the next steps.
Rules-based workflowThe sequence is known and conditions can be specified in advance. AI may handle an individual step.
AI agentThe next step depends on information gathered and the system must choose among authorised tools and actions. Evaluate the path as well as the final answer.

An example: preparing a customer support reply.

  1. Read the request and establish context

    The agent receives an in-scope request and identifies the required information. If essential details are missing, it drafts a question or passes the case to a person.

  2. Consult authorised sources only

    It may search approved documents and read order status after the requester and permissions are verified. It should not access every available record simply because it can.

  3. Prepare a reviewable proposal

    The draft includes useful references for the reviewer. Contradictory or unavailable information remains explicit instead of becoming an invented answer.

  4. Wait for the required approval

    In the first pilot, a person checks the proposal before sending. Refunds, order changes and commercial commitments remain outside the scope unless separately designed.

This is a design example, not a client result or a guarantee that an agent can handle every request reliably.

Controls are part of the product.

OWASP identifies prompt injection as an LLM risk: instructions inside external content can change system behaviour. Treat emails, pages and retrieved documents as untrusted data, not as authority to take action.

  • An allowlist of tools and actions, with reading separated from modification.
  • Limits on steps, execution time and cost, enforced when reached.
  • Human confirmation for specified actions and a route back to manual handling.
  • A record of sources and actions, with retention agreed for its contents.
  • Validation of tool inputs and outputs, including unexpected cases and malicious instructions.

Evaluate both the outcome and the actions taken.

Tests should include ordinary, ambiguous and out-of-scope requests, plus content that attempts to change instructions. Assess output quality, action correctness, access used, escalation and the human effort required to verify the work.

The first pilot uses a limited set of sources and tools. Agree in advance which errors block introduction and which require revision. Changes to the model, tools or documents may require another evaluation.

When an agent adds avoidable complexity.

We would not propose an agent just to transfer a field between systems, apply a formula or follow a known sequence. If access cannot be limited, actions cannot be checked or ownership is unclear, reduce autonomy or use an assistant that only prepares drafts.

FROM IDEAS TO A BRIEF

A template to work from.

First AI pilot brief

An editable Markdown template for your process, data, owners and go/no-go criteria.

Download the Markdown template

Practical questions

Can an AI agent work without supervision?

Decide autonomy action by action after evaluation. A first project can remain read-only or produce drafts. Actions affecting customers, data or money need proportionate controls and a named owner.

Can it use our internal documents?

We can assess retrieval from approved sources, checking freshness, permissions and quality. Retrieving a document does not guarantee a correct answer: reviewers need accessible references and tests on real questions.

How do you prevent excessive calls or endless execution?

Set limits on steps, duration and budget, along with stop conditions and alerts. The system executing tools should enforce the limits; they should not depend solely on text instructions to the model.

Which process should improve first?

Bring one specific task, the tools you use and the point where work gets stuck. We can define what is needed to assess a first project.

Let’s discuss your process