MazeTech · Agentic Systems Engineering · South Africa

Production AI agents for companies that need control, not experiments.

We design, build, and operate governed agentic systems: tool-using AI that acts inside defined contracts, approval gates, and telemetry. For South African enterprises and larger SMEs that are done with pilots that go nowhere.

Assessment output: a control gap map, a scoped memo, and one technical walkthrough. You keep the artifacts whether or not we build together. No logo wall, no invented metrics.

Fig. 01 · Controlled agentic system POLICY · LIFECYCLE · OWNERSHIP REQUEST TASK + CONTEXT AGENT PLAN · TOOLS · MEMORY SCOPED IDENTITY, LEAST PRIVILEGE TOOL A CONTRACT TOOL B CONTRACT TOOL C CONTRACT APPROVAL GATE HUMAN DECISION EXECUTION + FULL TRACE TELEMETRY BUS · EVERY ACTION RECORDED OBSERVABILITY DASHBOARD EVAL SUITE REGRESSION GATE AUDIT LOG
REQUEST
Task + context
AGENT
Plan · Tools · Memory · Scoped identity
TOOL CONTRACTS
Declared tools · Allowed operations · Data boundaries
APPROVAL GATE
Human decision · Denials logged as exceptions
EXECUTION + TRACE
Every action recorded
OBSERVABILITY · EVAL SUITE · AUDIT LOG
Telemetry bus, regression gate, audit trail
Market context Public research, cited as context. Not MazeTech performance.
33%

of enterprise software applications are projected to include agentic AI by 2028, up from under 1% in 2024.

Source: Gartner, December 2024
40%

of AI-related data breaches are projected to involve shadow AI by 2027, up from 7% in 2024.

Source: Gartner, 2025
TOP 10

OWASP maintains a Top 10 of LLM application risks. The failures cluster around agency, permissions, and missing evaluations, not model quality.

Source: OWASP LLM Top 10
Sec 01 Sheet 01 / 11 · Operating problem

Adoption is ahead of the guardrails.

Teams feel pressure to adopt AI, so they adopt it. Tools ship features, staff provision them, and shadow usage grows. The result is scattered experiments, duplicated spend, and data moving through tools nobody approved, with no single owner for systems that now touch operations.

Signature · Pilot stall

Pilots that end at the demo

Business owners ask for outcomes. The project stops at a notebook or a chatbot with no path to production. The budget line is spent; the workflow is unchanged.

Signature · Shadow usage

Consumer AI inside work data

Staff use consumer tools for work documents. The CISO inherits exposure without a control, and nobody can say what left the organisation.

Signature · Unbounded agents

Vendor agents without boundaries

Platform agents run with broad permissions and no approval step. The CIO discovers the blast radius after an incident, not before.

Signature · Procurement stall

No evidence to buy on

The CFO cannot see scope, controls, or exit criteria. Without artifacts the purchase cannot be defended, so it stalls.

Signature · Compliance exposure

Data movement without a map

Personal information moves between tools with no data map, no stated purpose, and no retention plan. Compliance and legal find out last.

Signature · No owner

Nobody accountable

When an agent misbehaves, issues bounce between the vendor and the team. There is no named owner, no runbook, no review cadence.

Who this is for One purchase, five reviewers
CEO / COO

Who is accountable when an agent acts?

One named owner per system, an escalation path, and a review cadence. Control is an operating feature, not a slide.

CIO / CTO

Does this fit our architecture?

Tool contracts describe any system we integrate. No platform swap, no lock-in. Lifecycle ownership is yours at handover.

CISO

What can the agent touch, and can we see it?

Dedicated identity, least privilege, and a full run trace for every action. Blast radius is designed, not discovered.

Compliance / Legal

Where does our data go, and who decides?

Data mapping, purpose limitation, retention, and an audit trail. POPIA-aware by design, with an exception register.

Procurement / CFO

What exactly are we buying?

Phased scope, named artifacts per phase, and go/no-go gates you control. Spend is tied to deliverables you keep.

Next step: request the assessment questionnaire. It is the same instrument we use in phase one, and you keep the output.

Sec 02 Sheet 02 / 11 · What MazeTech builds

Agents that work inside a system, not around it.

A governed agentic system has three fixed components. Everything else is configuration.

01

Tool contracts

Every tool an agent can touch is declared: name, purpose, allowed operations, data boundaries, rate limits, and the owner of the contract. The agent cannot call what is not in the contract.

02

Approval gates

Actions above a defined risk line pause for a named approver. The gate is part of the workflow, not an afterthought. Denials are logged as exceptions, not deleted.

03

Telemetry and evaluations

Every run is traced end to end. An eval suite guards each change. You can see what the agent did, why it did it, and whether the change improved anything.

Scope discipline

Every system also gets a named owner, a review cadence, and a retirement path. We do not build ChatGPT reskins, open-ended research projects, or agents that act on production data without a gate. If a workflow should not be automated, we say so in the assessment memo.

Sec 03 Sheet 03 / 11 · Failure modes

Most agent projects fail at the boundary, not the model.

Five failure modes show up repeatedly in production. Each has a known signature and a known control. We build the control in from the first commit.

Mode 01

Excessive agency

The agent can act anywhere. You learn about actions after the fact, from a customer or an auditor.

ControlScoped tool contracts and risk-ranked approval gates.

Mode 02

Broad permissions

The agent inherits the operator's identity, or a service account with everything. One credential now covers every system.

ControlDedicated agent identity with least privilege and per-tool scopes.

Mode 03

No evaluations

Every model update or prompt change is a blind release. Behaviour drifts between versions and no test catches it.

ControlAn eval suite in the delivery pipeline with a regression gate.

Mode 04

No telemetry

Nothing can be reconstructed after the fact. Incidents are investigated from memory and screenshots.

ControlRun traces, structured logs, and an audit trail on every action.

Mode 05

Unclear ownership

Issues bounce between vendor and team. Nobody owns the system, so nobody is accountable for its behaviour.

ControlA named system owner, review cadence, and retirement path.

The pattern

Every failure is a design decision.

None of these modes require heroic engineering to avoid. Each has a known control, applied at design time. That discipline is what MazeTech sells: not intelligence, but boundaries.

Sec 04 Sheet 04 / 11 · Production control model

The control stack.

Five layers, one rule: nothing executes without an owner, a contract, and a trace. Human oversight runs the length of the stack.

HUMAN OVERSIGHT · APPROVE · REVIEW · OWN
L1 Identity & least privilege
The agent runs on its own identity, scoped to the workflow it serves. It cannot borrow a person's access, and its scopes are reviewed on a schedule.
L2 Tool contracts
Declared tools, allowed operations, data boundaries, and a named contract owner. Anything outside the contract is a hard stop.
L3 Approval gates
Actions above the risk line pause for a named approver. Decisions, approvals, and denials are recorded as part of the workflow.
L4 Telemetry & tracing
Every run is traced end to end: inputs, tool calls, gate decisions, outputs. Any action can be reconstructed after the fact.
L5 Evaluations & lifecycle
Each system carries an eval suite, a release gate, a review cadence, and a retirement path. Changes ship only through the pipeline.
Policy rail

Policy, audit, and compliance run alongside the stack. High-impact decisions always involve a person. The stack is not a substitute for judgement; it is the structure that makes judgement accountable.

R1

An agent never holds broader access than its workflow requires.

R2

An action above the risk line never executes without an approval decision.

R3

A change never ships without passing the eval suite.

R4

A run never happens without a trace, and a trace never expires without a retention rule.

Sec 05 Sheet 05 / 11 · Use cases by operating area

Patterns we build per client.

These are reference patterns, not claims about delivered results. Each one ships with contracts, gates, telemetry, and named artifacts. We do not publish client names or metrics without their consent.

Finance operations

Invoice intake, reconciliation support, and variance flags. The agent prepares; a person disposes.

Workflow

  • Invoice and statement intake with extraction
  • Reconciliation matching against the ledger
  • Variance flags with supporting evidence

Controls in place

  • No payment execution. Approval required for any payment-adjacent action
  • Read-only access to ledgers via contract scopes
  • Every matched item carries a trace to its source document

Artifacts produced

  • Per-item processing trace
  • Exception register for disputed items
  • Reconciliation report for the finance owner
Owner · CFO / Finance manager

Customer operations

Structured response drafting, case triage, and SLA monitoring. Drafts are reviewed before any customer-facing send.

Workflow

  • Case triage against policy and history
  • Structured response drafting from approved templates
  • SLA and backlog monitoring with escalation flags

Controls in place

  • Customer-facing sends require a named reviewer
  • Template versioning; the agent cannot improvise commitments
  • Personal data handled per the data map and retention rules

Artifacts produced

  • Response trace with template version
  • Review and approval record
  • Escalation log for SLA flags
Owner · Operations manager

Procurement & supply

RFQ comparison, contract clause checks against policy, and vendor data assembly. Commitments stay human.

Workflow

  • RFQ response comparison against scored criteria
  • Contract clause review against policy checklists
  • Vendor information assembly with source citations

Controls in place

  • Any commitment or PO-adjacent action requires approval
  • Clause flags link back to the exact contract text
  • Supplier data access scoped to the evaluation window

Artifacts produced

  • Comparison memo with sources
  • Policy deviation list for legal review
  • Approval record for the buying decision
Owner · Procurement lead

Compliance & reporting

Evidence collection, control testing, and report assembly. The agent gathers; the accountable person certifies.

Workflow

  • Evidence collection against control inventories
  • Control testing support with sampled evidence
  • Report assembly from approved structures

Controls in place

  • No self-certification. Submissions require review by the accountable person
  • Evidence items carry source and collection time
  • Retention applied to collected evidence per policy

Artifacts produced

  • Evidence index with sources
  • Control test results and exceptions
  • Audit trail of the collection process
Owner · Compliance / Risk officer

IT operations

Change request triage, runbook drafting, and incident timeline assembly. Infrastructure changes keep the change process.

Workflow

  • Change request triage against impact classes
  • Runbook drafting from approved procedures
  • Incident timeline assembly from logs

Controls in place

  • Infrastructure changes require change approval; no direct execution
  • Runbook content locked to approved revisions
  • Access scoped to read-only log and ticket sources

Artifacts produced

  • Change record with impact assessment
  • Timeline trace for incident reviews
  • Runbook revision log
Owner · IT service manager

Next step: send us your workflow list. The assessment maps it against these patterns and tells you which one, if any, is worth piloting.

Sec 06 Sheet 06 / 11 · Governance, security, POPIA

Governance is a design input, not a slide.

Security and privacy controls are specified before the first line of agent code. Compliance teams review the design, not the aftermath.

01

Data mapping first

We map what the agent touches, where that data moves, and why, before any build. Purpose and minimality are POPIA obligations; they are also engineering inputs.

02

Purpose limitation in the contract

Tools receive data scopes, not blanket access. The contract states what the agent may read, transform, and produce, and what it may never do.

03

Retention and deletion

Logs, traces, and collected data follow a declared retention schedule with a deletion path. Nothing is kept because it is convenient.

04

Access and review

Named approvers for high-impact decisions. Access reviews run on a schedule, and scope creep is a reportable event.

05

Vendor processing

Any third-party model or tool is assessed as a processor: where data goes, what is retained, and what the contract says.

06

Exception register

Anything that deviates from the declared design is recorded, dated, and reviewed. Exceptions are managed, not hidden.

POPIA, stated plainly: the Protection of Personal Information Act governs how personal information is collected, used, and stored. We design to those obligations: documented purpose, minimal collection, and a trail of who accessed what. Where a certification is claimed, it is evidenced. We do not list certifications we do not hold, and we will not pretend a model vendor's terms replace your obligations.
Sec 07 Sheet 07 / 11 · Engagement model

Four phases. Each with scope, artifacts, and a gate you control.

No open-ended engagements. Every phase has fixed scope, named deliverables, and an explicit go/no-go decision before the next phase starts.

Phase 00

Assessment

1 to 2 weeks
  • Questionnaire and workflow inventory
  • Control gap map of current AI usage
  • Assessment memo with a recommendation
Gate: you review the memo before any build. You keep it either way.
Phase 01

Controlled pilot

3 to 6 weeks
  • One workflow, one tool contract
  • Gates on, telemetry on from day one
  • Pilot report against agreed criteria
Gate: measured against the criteria you approved. Go/no-go is yours.
Phase 02

Production build

Scoped per system
  • Full control stack and integrations
  • Eval suite wired into the pipeline
  • Operator training, runbook, handover
Gate: acceptance criteria from the pilot report, demonstrated in a walkthrough.
Phase 03

Operate & improve

Monthly cadence
  • Named owner with a review rhythm
  • Eval updates and drift checks
  • Exception register and roadmap reviews
Gate: renewal is by decision, not by default. Exit is clean.
Why it is structured this way

Procurement needs scope; operators need ownership; compliance needs a trail. The phase model gives all three. Spend is tied to deliverables you keep, and the riskiest phase, the pilot, is deliberately the smallest and most controlled one.

Sec 08 Sheet 08 / 11 · Proof model

We prove the work with artifacts, not adjectives.

Every engagement produces inspectable artifacts. If a claim cannot be backed by a document you can read, we do not make it.

Tool contract spec

Every tool, its scopes, boundaries, and owner, in writing.

Gate configuration

Risk lines, approvers, and decision records, versioned.

Eval suite and results

Test cases, pass criteria, and results per release.

Run traces

End-to-end records of what each agent did and why.

Architecture decision records

Why each control exists, dated and attributed.

Exception register

Every deviation, its date, and its review.

Assessment memo

The control gap map you keep before any build.

Handover runbook

Operations, escalation, and retirement in one document.

Design references NIST AI RMF · risk management OWASP LLM Top 10 · application risks ISO/IEC 42001 · AI management systems POPIA · personal information

We build against these references and can show where each control maps. Alignment is documented in the artifacts. Certification is only claimed with evidence.

Evidence policy

We do not publish invented metrics, client logos, or case studies. If you want proof, we give you artifacts you can inspect: specifications, configurations, traces, and evaluation reports. Before you buy, we walk you through the assessment memo, the control map, a live run trace, and the eval suite. You see the evidence, then you decide.

Sec 09 Sheet 09 / 11 · Objections

Fair questions, direct answers.

If you are defending this purchase to a board, a risk committee, or a CFO, these are the questions you will hear.

Most pilots fail at governance and scope, not model quality. The assessment starts from your existing attempts and reviews their artifacts. If the blocker was data or workflow, we say so and scope accordingly.

The assessment tells you honestly. Sometimes the correct answer is to fix the data or the workflow first, and we will write that in the memo. The memo is yours whether or not we build together.

The deliverable of phase one is a memo with a control gap map you can keep and circulate. We are paid for artifacts, and you keep every one of them. There is nothing to sign before you have seen them.

Data map, purpose limitation, retention schedule, processor review, and an audit trail. POPIA-aware by design, documented before build. Your data stays inside the boundaries the contract states.

That is the point of a controlled pilot: small scope, gated actions, and criteria you approved in advance. If the criteria are not met, you have an exception record and a clean exit, not a sunk project.

Tool contracts describe any system, from an ERP to a shared spreadsheet. The contracts are the integration. We do not require a platform swap, and the assessment memo identifies integration points before any build.

Each phase has fixed scope and named artifacts. Spend is tied to deliverables, and every phase ends in a go/no-go gate that you control. The assessment memo defines the scope of the pilot in writing before it starts.

You do. Lifecycle ownership is built into the engagement: a named owner, a runbook, a review cadence, and an operate phase that continues until your team is ready to run it alone.

Sec 10 Sheet 10 / 11 · Start

Start with the assessment.

One questionnaire, one control gap map, one walkthrough. You keep the artifacts whether or not we build together.

  1. 1You send a one-paragraph brief: the workflow, the pain, the stakeholders.
  2. 2We return the assessment questionnaire within two working days.
  3. 3We deliver the memo and walk you through the control gap map.
Assessment request

Subject line: Agentic Systems Assessment. Include the workflow you have in mind and the buyers who will review the decision.

Send the brief