Projects

Case study

AI-enabled sales system

Building a lead engine that could follow up with every prospect

~50,000 AI calls in a peak month · ~400 autonomously booked meetings per month · supporting ~$1M/month in sustained cash collection

A US business selling high-ticket packages had outgrown the way it handled leads.

Marketing was working. Leads were arriving from marketplaces, paid social, the website, inbound calls and social channels throughout the day. The problem was what happened next.

A four-person offshore setter team manually worked a queue during US Eastern business hours. At higher volumes, a new prospect might wait hours, sometimes until the following day, before somebody called them.

Follow-up depended heavily on the individual setter. CRM hygiene was inconsistent. Some prospects received repeated attention while others were missed entirely. As lead volume increased, adding more setters was not an attractive way to scale.

I progressively rebuilt that layer around AI voice, CRM orchestration and automated follow-up, creating a system that could contact, qualify and follow up with each prospect until they either booked a meeting, explicitly declined or moved into another defined state.

At maturity, the system was capable of roughly 50,000 calls in a peak month and consistently booked around 400 meetings per month into human closers’ calendars. It supported a sales operation that sustained roughly $1 million per month in cash collection for a period of years.

That growth was not caused by installing an AI caller. Marketing strategy drove substantially greater lead volume over time; the system was the infrastructure that allowed the business to handle that volume consistently without scaling the setter layer alongside it.

At a glance

lead-intake channels
10
opportunities in the core CRM pipeline
~45K
calls in a peak month
~50K
meetings booked autonomously per month
~400

The audited production environment contained roughly 45,000 opportunities, 41,000 unique contacts, 83 published GoHighLevel workflows, 22 active Make scenarios and 47 Vapi assistants, although only a subset of the assistants were live at any one time.

Where we started

Before the voice system, four offshore setters handled first contact and follow-up manually. They worked standard US Eastern hours. A new enquiry entered the queue and waited until someone became available to call it.

At low volume, that model could work. At higher volume, it produced several structural problems.

Speed depended on capacity

Leads arrived outside working hours and faster than the team could consistently process them. Some waited hours before first contact. Others did not get called until the next day. For a high-intent lead who had just submitted an enquiry, that delay mattered.

Follow-up depended on the person

There was no sufficiently reliable reporting to prove how many attempts each lead actually received or whether the prescribed cadence was being followed. The practical reality was that it was not consistent. Some leads received significantly more attention than others.

CRM state could not be trusted

Manual setters did not maintain records consistently. Contacts were missed, statuses were ambiguous and the CRM did not always provide a reliable representation of where each prospect actually sat in the sales process.

Scale meant adding people

If lead volume doubled, the obvious answer was more setters. That introduced more hiring, training, management and variable execution, while still failing to guarantee that every lead would receive the same treatment.

The objective was therefore not simply:

Make an AI phone call.

It was:

Create a lead-handling system where every prospect could receive immediate, persistent and consistent attention regardless of volume.

The setting layer, before and after

4 setters

  • Business-hour queue
  • Variable contact delay
  • Manual cadence
  • Inconsistent CRM updates

AI setting layer

  • Near-immediate response
  • Concurrent capacity
  • Defined state and cadence
  • Structured CRM output

What I took over

I owned the technical system end to end. That included the architecture, CRM design, workflow logic, integrations, voice-agent architecture, prompts, troubleshooting, documentation, optimisation and the decisions around how leads should move through the system.

As the platform became larger, I had a developer assisting with maintenance, upgrades and implementation. I remained responsible for the architecture and technical direction.

One of the first problems was not AI at all. It was inheritance.

The previous implementation had become difficult to operate. Documentation was effectively nonexistent. Workflows, fields and integrations were ambiguously named, inconsistently structured and in some cases simply misspelled. As dependencies accumulated, understanding the effect of changing one component became increasingly difficult.

That experience heavily influenced how I approached not just this system, but later builds:

Documentation and system design are part of the infrastructure. They are not administrative work added after the infrastructure exists.

Design principles

A few principles eventually shaped the architecture.

One system should hold the truth

The CRM became the state machine. Make moved information between services. Vapi handled the conversation. But neither of those systems independently decided what happened to the lead next. They wrote facts back into the CRM. The CRM then acted on those facts.

This made the operational state visible to the business rather than burying it inside integration middleware. In production, Make did not directly move a lead through pipeline stages; it wrote outcomes into fields and CRM workflows reacted to them.

State should drive behaviour

A brand-new enquiry should not receive the same conversation as someone who:

  • asked for a callback;
  • expressed interest but did not book;
  • missed a scheduled meeting;
  • stopped answering;
  • had previously said no.

Each state therefore had its own handling and, where useful, its own specialist voice agent.

AI should not own every decision

Conversational ambiguity is useful territory for an LLM. System state is not. The architecture increasingly separated three kinds of decision.

AI
  • Conversation
  • Extracting intent
  • Understanding natural-language dates
  • Answering off-script questions
  • Summarising calls
Rules
  • Pipeline state
  • Routing
  • Cadence
  • Exit conditions
  • Closer assignment
  • Alerts
  • Whether an appointment write succeeded
Human
  • High-value sales conversations
  • Edge cases
  • Customer support
  • Failed bookings
  • Relationship-sensitive scenarios

Failures should become visible

A failed automation should not silently disappear. Unrecognised enquiry sources, failed data extraction and appointment-write errors were designed to stop or raise an alert rather than allow bad data to propagate through the system.

System architecture

At the centre were three platforms:

GoHighLevel
State, pipeline, calendars, SMS/email and workflow logic.
Make
Integration and orchestration between services.
Vapi
Live AI voice conversations.

Around them sat the acquisition sources, telephony, models, speech services, knowledge layer, newsletter infrastructure, website and reporting systems. The production stack ultimately touched 16 platforms.

  • Rules deterministic logic
  • AI a language model
  • Human a person
  1. Lead sourcesMarketplaces · Meta · Website · Phone · Social
  2. Intake Rules
  3. Extraction from messy sources AI
  4. Normalisation and dedupe Rules
  5. GoHighLevel: CRM and state Rules
  6. SMS / email Rules
    1. Make Rules
    2. Vapi voice agent OpenAI · voice · speech-to-text AI
    3. Live call AI
    4. Availability and knowledge lookup AI
    5. Post-call result AI
    Direct booking Rules
  7. GoHighLevel: pipeline state Rules
  8. Appointment booked Rules
  9. Human closer Human
Every route out of the CRM leads back into it. The CRM decides what happens next.

Ten ways into one system

The system received prospects through ten intake channels. These included high-volume marketplace enquiries, Meta lead forms, paid-social landing pages, the main website, inbound telephone calls, social DMs and several manually triaged partner sources.

The important architectural decision was that these did not become ten separate sales systems. They converged into one model.

One discriminator

A custom lead-source field encoded both the relevant brand and acquisition path. Downstream logic could therefore ask:

Where did this prospect come from?

and derive:

  • which brand they belonged to;
  • which voice agent should call;
  • which calendar should be used;
  • which nurture track they should enter;
  • which cadence variant should run.

That avoided duplicating entire CRM pipelines for every source.

The system had originally used separate 19-stage pipelines for brands. These were eventually collapsed into one shared 25-stage pipeline, making a new brand primarily a data/routing change rather than another duplicated architecture.

Store variation as data, not duplicated structure.

A new brand should require a new value and a few routing branches, not a cloned CRM architecture.

Making messy intake usable

Not every lead arrived cleanly. Marketplace leads, for example, often arrived as emails rather than structured webhook payloads. A typical flow was:

  1. Marketplace enquiry
  2. Email
  3. RulesMail rule
  4. RulesMake webhook
  5. AILLM extraction
  6. RulesSource lookup
  7. RulesDeduplication
  8. CRM

The LLM extracted fields such as:

  • name;
  • email;
  • phone;
  • comments;
  • source identifier;
  • URL.

The source identifier was then looked up to determine which brand owned the enquiry. If the system could not identify the source, it did not guess. It stopped and alerted the team. Likewise, parsing failures stopped the scenario before bad information reached the CRM.

Missing phone numbers were routed into a separate state and nurtured through email rather than accidentally entering the dial queue.

This mattered because AI orchestration is only useful when the state beneath it is reliable.

The CRM was the state machine

The core pipeline eventually contained 25 stages across four broad zones.

Lead states, by who owns them

  1. Intake

    Automation
    • New, with phone
    • New, without phone

    New leads entered with or without a valid phone number. Automation owned this stage.

  2. AI-worked

    AI
    • Callback requested
    • Interested but not booked
    • Unsuccessful AI follow-up
    • Meeting no-show
    • Short-term no-answer
    • Not interested
    • Text-only
    • Undecided

    The voice and messaging system owned most of these states.

  3. Booked

    Rules
    • Meeting booked
    • Reminders and closer alert

    Once a meeting existed, the lead transitioned toward human ownership.

  4. Post-meeting

    Human
    • Human closer
    • Sale or loss

    After the closer met the prospect, AI outbound calling stopped. The human sales process took over.

The data contract

About a dozen CRM fields formed the interface between GoHighLevel, Make and Vapi. Important examples included:

Lead source
Brand and acquisition source.
Call outcome
The structured result returned after each call.
Consecutive misses
How many unanswered attempts had occurred in sequence.
Contact status
Including whether booking succeeded or failed.
Dial state
Whether Vapi accepted the outbound request.
Call record
Summary, transcript, call type, number and timezone.

This is where the project becomes more interesting technically.

The voice model was not the system. It was one component producing structured inputs into a larger state machine.

The AI voice layer

Over the life of the system, 47 Vapi assistants existed across production, testing, previous versions and future iterations. Only around ten were typically live simultaneously.

47 assistants created~10 live at any one time

The system evolved through specialised assistants for different contexts. These included:

  • First contact, per brand
  • Paid-social first call
  • Paid-social follow-up
  • Callbacks
  • Interested but not booked
  • General follow-up
  • No-show rebooking
  • Missed-call callbacks
  • Reactivation
  • Inbound calls
  • Newsletter-sourced leads

The reason for specialisation was simple:

Someone asking for a callback should not hear the same opener as a cold lead.

The system’s state gave the agent contextual knowledge before the call began.

Prompt architecture

The voice agents used long system prompts, typically around 5,000–16,000 characters. The structure became reasonably standardised:

  1. Role and persona
  2. Delivery rules
  3. Lead context
  4. Conversation flow
  5. Scheduling
  6. FAQ and objections
  7. Guardrails

The later scripts moved away from rigid pitching toward a more discovery-led interaction: understanding the prospect’s situation, explaining the model, reframing objections and then moving toward scheduling. Some prompts contained roughly 20 scripted objections.

We also spent considerable time on something much less impressive in a demo but critical in production: how the agent sounded.

Pacing, interruption handling, response latency, speech models, filler language and turn-taking all materially changed whether a conversation felt usable.

At the time, voice models were moving extremely quickly. Configurations that had taken considerable effort to make acceptable could become obsolete months later as latency and model quality improved.

What one lead actually experienced

One of the easiest ways to understand the architecture is to follow a marketplace enquiry through it.

  1. 0 sec

    Enquiry submitted

    A prospect submits an enquiry through a marketplace.

  2. + seconds

    Parsed and deduped AI

    The notification email reaches a mailhook. An LLM extracts the contact data and source identifier. The system determines the correct brand and checks the CRM for an existing record.

  3. + seconds

    CRM state set Rules

    The CRM creates or updates the opportunity, stamps its source and:

    • sends the relevant welcome SMS;
    • enters it into the appropriate nurture track;
    • enrols it into AI calling.
  4. + ~15 sec

    Outbound call initiated Rules

    After a deliberately small delay, the dialer passes the prospect to the correct voice agent using a local number. In some flows, we deliberately sent SMS first or introduced a short pause rather than creating an unnaturally instantaneous phone call.

  5. During the call

    Live qualification and availability AI

    The agent:

    • confirms the enquiry;
    • explains the offer;
    • qualifies;
    • handles objections;
    • identifies the prospect’s location/timezone;
    • queries live calendar availability;
    • agrees a meeting time.
  6. + ~40 sec after the call

    Transcript processed, appointment written AI

    The transcript processor:

    • classifies the result;
    • saves the call;
    • extracts the agreed appointment;
    • writes it into the calendar;
    • moves the prospect into the appropriate pipeline state.
  7. T−24 hr

    Prospect and closer reminders Rules

    The system sends prospect reminders and alerts the assigned closer. If the prospect no-shows, separate email, SMS and AI voice recovery processes can begin.

  8. Meeting

    A human takes over Human

The hardest voice problem: trust the tool, not the model

One recurring failure mode came from scheduling. The agent was supposed to check live calendar availability before agreeing to a meeting.

But if the availability tool failed, or if the conversation diverged in an unexpected way, the model could still behave conversationally and accept a proposed time. To the prospect, the call sounded successful. Operationally, it wasn’t.

The downstream workflow would later attempt to write the appointment and discover that the agreed time did not exist.

That created a useful design lesson:

A conversationally plausible answer is not the same thing as a valid business action.

The system eventually treated booking as a two-part process. During the call: check and agree. After the call: validate and write.

The model checked live availability while speaking with the lead. Once the call ended, the transcript was analysed again, the agreed time was extracted in the prospect’s timezone and the appointment was written into the correct calendar.

That kept slower CRM writes out of the live conversation and made failed appointment writes detectable rather than silently lost.

During the callCheck and agree

  1. Lead on the call
  2. Asks city and timezone AI
  3. Requests availability AI
  4. Make parses the date and time AI
  5. GoHighLevel free-slots API Rules
  6. Available: agree the time AI
    Not available: propose alternatives AI
  7. End call

After the callValidate and write

  1. Transcript analysis AI
  2. Extract the agreed timestamp AI
  3. Write the appointment Rules
  4. Written: booked stage Rules
    Failed: human alert Human
The call agrees a time. Only the calendar write makes it a booking.

Follow-up was a system, not a retry button

A lead who failed to book was not simply called again forever. Pipeline state determined the next strategy.

Different stages had different:

  • scripts;
  • retry windows;
  • call caps;
  • cooling-off periods;
  • voice agents.

The main stages used 8–16 attempts across roughly 7–10 days, depending on the state and brand. A typical no-answer pattern might begin aggressively before backing off:

A typical no-answer pattern

  1. 2 min
  2. 2 min
  3. 2 hr
  4. next day
  5. backs off

The point was to capture high-intent demand while it was fresh without blindly hammering every lead indefinitely.

A state-driven loop

  1. Call AI
  2. Outcome AI
  • Booked Stop AI, send reminders, hand to the closer
  • Interested Dedicated follow-up state
  • Callback Callback state
  • Unclear AI follow-up state
  • No answer Add to the miss count, retry
  • Not interested Exit the active cadence

A universal exit process removed leads from calling sequences when they:

  • booked;
  • entered closing;
  • won;
  • no-showed into another workflow;
  • explicitly declined;
  • called the business themselves.

This prevented the particularly bad automation failure of calling someone with a sales pitch after they had already booked.

Number reputation became a system constraint

One of the more important lessons was that the limiting factor was not always AI capability. It was telecommunications infrastructure.

Long-term aggressive calling eventually caused carriers to flag outbound numbers as spam. A long-term no-answer campaign was therefore shut down, and voicemail drops that had been built were left disabled.

The architecture adapted through:

  • local numbers;
  • a pool of 65 outbound numbers;
  • drip throttling;
  • per-stage call limits;
  • reduced long-tail calling;
  • greater reliance on email reactivation.

This is a good example of why an AI system cannot be designed in isolation from the environment it operates in.

The human handoff was deliberate

The objective was never to remove humans from the entire sale. The product was high ticket and consultative. A human closer remained more valuable once a prospect had committed enough attention to book a meeting.

So the system’s boundary became:

AI owns
  • Initial contact
  • Qualification
  • Routine objections
  • Follow-up
  • Appointment setting
  • No-show recovery
  • Selected reactivation
Rules own
  • State
  • Routing
  • Timing
  • Calendars
  • Retries
  • Exits
  • Alerts
Humans own
  • Closing calls
  • Complex edge cases
  • Relationship-sensitive scenarios
  • Support
  • Failed bookings requiring intervention
  • The sales process after the meeting

The setter layer gradually reduced from four people. Three poorly performing setters left the process, while one moved into a full-time closer role. Setters were still occasionally used for edge cases that the system could not appropriately handle.

That was the point of the architecture:

Move humans up the value chain rather than forcing them to perform repetitive work the system could handle more consistently.

Reliability mattered more than the demo

A demo usually shows:

Lead arrives → AI calls → meeting booked.

Production has to answer:

What happens when something fails at every step?

The system eventually included patterns such as:

  • Retries

    CRM writes and lookups retried automatically rather than dropping the run on transient errors.

  • Fail closed on bad data

    Unknown source identifiers or extraction errors stopped processing instead of guessing.

  • Booking alerts

    If a prospect verbally agreed to a meeting but the calendar write failed, an internal alert was generated for manual recovery.

  • Test isolation

    Known test records could be diverted before triggering real calls.

  • Central state

    Middleware reported facts; the CRM made business-state decisions.

The production system deliberately used incomplete-execution queues and retry handlers so failed integrations could be replayed rather than silently lost.

The system measured itself

Every AI call generated more than a transcript.

Post-call analysis

  • Call outcome
  • Booked time
  • Main objection
  • Whether the prospect appeared to detect the AI
  • Their reaction if so
  • Where the agent performed poorly
  • Call summary
  • Success score

Analytics API, per call

  • Duration, result and agent
  • Objection and recording
  • Total cost
  • Speech-to-text cost
  • LLM cost
  • Voice cost
  • Platform cost
  • Carrier cost

The post-call analysis created a feedback loop for prompt iteration. A custom analytics API also received per-call information, including a full cost breakdown.

The purpose was not just reporting. It allowed us to ask:

  • Which agent is expensive?
  • Which script fails most often?
  • What are prospects objecting to?
  • Where does the AI break?
  • What does a booked meeting cost?

The CRM timeline then gave the closer a consolidated history containing AI call summaries and transcripts alongside SMS and email activity.

The feedback loop

  1. Call AI
  2. Transcript AI
  3. Structured analysis AI
    • Outcome
    • Objection
    • AI detection
    • Weak point
    • Summary
  4. CRM and analytics Rules
  5. Prompt or flow change Human
  6. Next version
The call was not only an action. It was data for improving the next call.

It was never “set and forget”

This was probably the most important operational lesson. The system evolved constantly.

Voice models improved. Speech latency changed. Prompting approaches improved. Business offers changed. Brands were added. Marketing channels changed. Call reputation changed. The CRM evolved.

What made sense in one Make scenario six months earlier could often be significantly simplified when a platform later introduced native functionality.

The technical audit itself found a system with 83 live CRM workflows from 158 built, 22 active Make scenarios in an account containing 163, and around 10 active voice assistants among 47 total versions.

Production footprint

CRM workflows
83live of 158 built
Make scenarios
22active of 163
Voice assistants
~10live of 47 versions
Outbound numbers
65
Integrated platforms
16
Intake channels
10

Live or activeBuilt over the system’s life

Some of that is the inevitable archaeology of a long-running production system. But it also reinforces something I now consider fundamental:

Automation needs ownership.

A system this interconnected does not become “finished” because the first version works. It needs monitoring, maintenance, experimentation, documentation and periodic simplification.

What went wrong

  1. We inherited an undocumented system

    The previous developer left virtually no usable documentation. Naming was inconsistent and often ambiguous. Even spelling errors became part of the implicit contract between workflows.

    This made troubleshooting unnecessarily slow and changes unnecessarily risky.

    What changed

    Architecture documentation, naming discipline and future maintainability became explicit requirements.

    Later, giving AI coding tools direct access to APIs and the surrounding system also dramatically improved the ability to trace and understand dependencies.

  2. The agent could sound right while being operationally wrong

    If the availability tool failed, an LLM could still conversationally accept a time. The prospect believed they had booked. The system had not.

    What changed

    Booking was treated as a controlled transaction rather than trusting conversational output. Failures became explicit states with alerts.

  3. Timezone handling caused bad calling windows

    An early configuration could result in East Coast prospects receiving calls at inappropriate local hours.

    What changed

    Calling was standardised around a safer operating window, while agents captured timezone context for scheduling. The audit still identified local-time scheduling as an area for further improvement.

  4. Carrier reputation limited theoretical scale

    More calling was not automatically better. Long-term campaigns damaged number reputation and reduced deliverability.

    What changed

    Long-term calling was reduced, voicemail drops stayed disabled and more reactivation moved to email.

  5. Old integrations accumulated debt

    The technology landscape moved quickly. Some Make scenarios that had been sensible when built could later be replaced or dramatically reduced as platforms developed better APIs or native integrations.

    What changed

    Regular architectural review became part of maintaining the system rather than assuming existing workflows remained optimal.

  6. Dependencies can fail invisibly

    At one point, a knowledge-base backend was disabled and dozens of agents lost their fallback capability without producing an obvious system-wide error. The audit identified dependency health checks as a missing safeguard.

    What changed, conceptually

    A dependency being reachable needs to be monitored separately from whether the overall workflow appears “on”.

Results

The business was collecting roughly $200,000 per month before the system matured. As marketing increased lead volume, the AI-enabled sales system became the infrastructure that allowed the business to handle that growth without scaling the setter team alongside it.

At maturity, the system supported:

AI calls in a peak month
~50K
meetings booked autonomously each month
~400
sustained cash collection across the sales operation
~$1M/mo

The system did not create that revenue in isolation. It supported the scale required to handle a much larger volume of inbound demand consistently.

Monthly cash collection

Before the system matured
~$200K

Acquisition and automation channels added progressively

Mature operating period
~$1M
Revenue growth resulted from broader commercial and marketing activity.The lead engine supported the volume required to operate at this scale.

Every lead could now move through a defined process: near-immediate first contact, persistent follow-up, structured CRM state, automated booking and a clear handoff to a human closer.

The result was not simply more calling capacity. It was a sales infrastructure that could absorb substantially more demand without allowing lead management to become the bottleneck.

What I learned

Three lessons have carried into almost everything I have built since.

Design for the system you will have to maintain

The quickest implementation is not necessarily the quickest system. Ambiguous names, undocumented dependencies and duplicated logic create a tax on every future change.

The more capable AI coding becomes, the more tempting it is to create quickly. That makes architecture and documentation more important, not less.

Separate intelligence from authority

An LLM can be excellent at interpretation without being the right mechanism for making every decision.

A model can understand

“Thursday sometime around two.”

A deterministic system should still decide

Whether two o’clock is actually available and what state transition follows.

Giving AI the parts it is good at while constraining authority elsewhere made the system much more reliable.

Automation is an operation, not an installation

This system required:

  • Monitoring
  • Prompt iteration
  • Testing
  • Carrier management
  • Workflow maintenance
  • Reporting
  • Cleanup
  • Documentation
  • Model upgrades
  • Process changes

The idea that a business can “turn on AI” and leave it untouched misunderstands the nature of production automation. A system that touches customers and revenue needs an owner.

What I would build differently today

If I were rebuilding the system from zero today, I would simplify it substantially.

  • Reduce platform dependence

    Several functions that historically required Make or complicated CRM workflows could now be implemented more cleanly in custom services or newer native integrations.

  • Make contracts explicit

    Instead of state depending on free-text fields and exact strings, I would use enumerated states and stable identifiers wherever possible.

  • Design local-time calling from the beginning

    Each lead’s timezone would become part of scheduling and cadence immediately rather than being captured but only partially operationalised.

  • Add dependency health monitoring

    Every external tool would have explicit health monitoring: voice, calendar, knowledge base, CRM and model endpoint.

  • Treat observability as a first-class system

    Errors, costs, latency, conversions, call quality and dependency failures would live in one consolidated observability layer rather than being distributed across tools.

  • Prune continuously

    Production systems accumulate old versions. Retirement should be part of the development lifecycle rather than a housekeeping exercise undertaken later.

Full technology stack 14 tools
LayerTechnologyPurpose
CRM / stateGoHighLevelPipeline, workflows, calendars, SMS/email
IntegrationMakeData movement, utility AI calls, API orchestration
VoiceVapiOutbound/inbound AI calling
TelephonyTwilioNumbers, carrier layer, SMS
LLMOpenAIConversation, extraction, date/time interpretation
VoiceElevenLabsText-to-speech
TranscriptionDeepgramSpeech-to-text
KnowledgeVoiceflow / ClaudeOff-script knowledge lookup
AcquisitionMetaLead forms, paid social, conversion signals
WebWordPress / ElementorLanding pages and booking
SocialManyChatDM/comment automation
NurtureBeehiivLong-term email nurture
ReportingGoogle Cloud RunCall analytics and cost ingestion
OperationsClickUpDownstream delivery operations

The phone call was the visible part.

The interesting engineering happened around it. A production AI sales system needed to know:

  • who the prospect was;
  • why they had entered;
  • which brand they belonged to;
  • what had already happened;
  • which conversation was appropriate now;
  • when to call again;
  • when to stop;
  • whether the calendar was genuinely available;
  • whether the booking actually succeeded;
  • when to bring in a person;
  • and what to do when any dependency failed.

That is the difference between an AI demo and an operating system.

The model handled the conversation. The surrounding architecture made it useful to the business.