Case study · AI
AI agents for clinical operations support
Vertex wanted agents that reduce repetitive support load without becoming an unsupervised black box. We shipped a narrow production lane first — then expanded coverage with evaluation.
Vertex Health12 min read
- Client
- Vertex Health
- Industry
- Healthcare operations
- Category
- AI
- Engagement
- Agent program
- Duration
- 14 weeks
- Published
- Mar 4, 2026
A production agent lane that reduces routine load — without hiding exceptions or bypassing clinical ownership.
Context
Vertex Health’s clinical operations teams were absorbing a rising volume of repetitive triage: classification, drafting responses and routing work to the right queue. Staff were skilled — and exhausted by work that did not need clinical judgment every time.
Leadership had already seen chatbot demos. None of them respected permissions cleanly. None of them escalated with enough context. The risk of an unsupervised system in a clinical setting was obvious.
The challenge
The organization needed relief that would survive compliance review. That meant agents with narrow tools, explicit data access, measurable quality and a hard stop when confidence dropped.
The political risk was as real as the technical one. If the first launch felt like a black box, operators would refuse to trust later lanes — regardless of model quality.
Approach
How we worked the problem
01
Narrow the first lane
We refused a broad chatbot brief. The first production scope was classify, draft and route — nothing that could silently act on high-risk steps.
02
Define evaluation first
Gold examples and failure modes were agreed before launch so quality was measurable from week one, not debated after incidents.
03
Enforce human handover
Low-confidence cases paused for operator review. Tools and data access were explicit, audited and limited to the lane.
04
Expand only when stable
Coverage grew after the first lane proved operable — not after a demo looked impressive in a slide deck.
Approach
We designed the program as a sequence of governed lanes, not a single omnipotent assistant. The first lane classified incoming work, drafted a suggested response and routed it. It never closed a case on its own.
Evaluation came before go-live. Clinical and operations leads agreed on gold examples, unacceptable failure modes and the confidence threshold that forced human review.
Access control was treated as product design. The agent could only call approved tools. Every action was logged. Exceptions landed in a queue operators already knew how to work.
What we delivered
The production lane reduced repetitive triage while keeping high-risk steps under human ownership. Operators received drafts with context, not opaque recommendations. Reviewers could accept, edit or escalate without leaving their operating surface.
We also left Vertex with the operating system around the agent: evaluation cadence, observability for handoffs, and a playbook for expanding into the next lane only after quality held.
Outcome
Routine work moved faster without hiding exceptions. Operators kept ownership of high-risk steps. Leadership could show auditors how the system behaved — including when it chose not to act.
The program expanded after the first lane was stable. That sequencing mattered more than any single model upgrade. Trust was earned in production, not promised in a workshop.
Deliverables
- Production agent lane for triage and routing
- Permission and tool-access model for clinical data
- Human review queue for low-confidence cases
- Evaluation set with agreed quality thresholds
- Operator playbooks and escalation rules
- Observability for agent decisions and handoffs
- Focus
- Support agents
- Mode
- Human review
- Lane
- Production
- Engagement
- 14 weeks
Related service: AI agent development