A US business selling high-ticket packages had outgrown the way it handled leads.
Marketing was working. Leads were arriving from marketplaces, paid social, the website, inbound calls and social channels throughout the day. The problem was what happened next.
A four-person offshore setter team manually worked a queue during US Eastern business hours. At higher volumes, a new prospect might wait hours, sometimes until the following day, before somebody called them.
Follow-up depended heavily on the individual setter. CRM hygiene was inconsistent. Some prospects received repeated attention while others were missed entirely. As lead volume increased, adding more setters was not an attractive way to scale.
I progressively rebuilt that layer around AI voice, CRM orchestration and automated follow-up, creating a system that could contact, qualify and follow up with each prospect until they either booked a meeting, explicitly declined or moved into another defined state.
At maturity, the system was capable of roughly 50,000 calls in a peak month and consistently booked around 400 meetings per month into human closers’ calendars. It supported a sales operation that sustained roughly $1 million per month in cash collection for a period of years.
That growth was not caused by installing an AI caller. Marketing strategy drove substantially greater lead volume over time; the system was the infrastructure that allowed the business to handle that volume consistently without scaling the setter layer alongside it.
At a glance
- lead-intake channels
- 10
- opportunities in the core CRM pipeline
- ~45K
- calls in a peak month
- ~50K
- meetings booked autonomously per month
- ~400
The audited production environment contained roughly 45,000 opportunities, 41,000 unique contacts, 83 published GoHighLevel workflows, 22 active Make scenarios and 47 Vapi assistants, although only a subset of the assistants were live at any one time.
Where we started
Before the voice system, four offshore setters handled first contact and follow-up manually. They worked standard US Eastern hours. A new enquiry entered the queue and waited until someone became available to call it.
At low volume, that model could work. At higher volume, it produced several structural problems.
Speed depended on capacity
Leads arrived outside working hours and faster than the team could consistently process them. Some waited hours before first contact. Others did not get called until the next day. For a high-intent lead who had just submitted an enquiry, that delay mattered.
Follow-up depended on the person
There was no sufficiently reliable reporting to prove how many attempts each lead actually received or whether the prescribed cadence was being followed. The practical reality was that it was not consistent. Some leads received significantly more attention than others.
CRM state could not be trusted
Manual setters did not maintain records consistently. Contacts were missed, statuses were ambiguous and the CRM did not always provide a reliable representation of where each prospect actually sat in the sales process.
Scale meant adding people
If lead volume doubled, the obvious answer was more setters. That introduced more hiring, training, management and variable execution, while still failing to guarantee that every lead would receive the same treatment.
The objective was therefore not simply:
Make an AI phone call.
It was:
Create a lead-handling system where every prospect could receive immediate, persistent and consistent attention regardless of volume.
The setting layer, before and after
4 setters
- Business-hour queue
- Variable contact delay
- Manual cadence
- Inconsistent CRM updates
AI setting layer
- Near-immediate response
- Concurrent capacity
- Defined state and cadence
- Structured CRM output
What I took over
I owned the technical system end to end. That included the architecture, CRM design, workflow logic, integrations, voice-agent architecture, prompts, troubleshooting, documentation, optimisation and the decisions around how leads should move through the system.
As the platform became larger, I had a developer assisting with maintenance, upgrades and implementation. I remained responsible for the architecture and technical direction.
One of the first problems was not AI at all. It was inheritance.
The previous implementation had become difficult to operate. Documentation was effectively nonexistent. Workflows, fields and integrations were ambiguously named, inconsistently structured and in some cases simply misspelled. As dependencies accumulated, understanding the effect of changing one component became increasingly difficult.
That experience heavily influenced how I approached not just this system, but later builds:
Documentation and system design are part of the infrastructure. They are not administrative work added after the infrastructure exists.
Design principles
A few principles eventually shaped the architecture.
One system should hold the truth
The CRM became the state machine. Make moved information between services. Vapi handled the conversation. But neither of those systems independently decided what happened to the lead next. They wrote facts back into the CRM. The CRM then acted on those facts.
This made the operational state visible to the business rather than burying it inside integration middleware. In production, Make did not directly move a lead through pipeline stages; it wrote outcomes into fields and CRM workflows reacted to them.
State should drive behaviour
A brand-new enquiry should not receive the same conversation as someone who:
- asked for a callback;
- expressed interest but did not book;
- missed a scheduled meeting;
- stopped answering;
- had previously said no.
Each state therefore had its own handling and, where useful, its own specialist voice agent.
AI should not own every decision
Conversational ambiguity is useful territory for an LLM. System state is not. The architecture increasingly separated three kinds of decision.
- Conversation
- Extracting intent
- Understanding natural-language dates
- Answering off-script questions
- Summarising calls
- Pipeline state
- Routing
- Cadence
- Exit conditions
- Closer assignment
- Alerts
- Whether an appointment write succeeded
- High-value sales conversations
- Edge cases
- Customer support
- Failed bookings
- Relationship-sensitive scenarios
Failures should become visible
A failed automation should not silently disappear. Unrecognised enquiry sources, failed data extraction and appointment-write errors were designed to stop or raise an alert rather than allow bad data to propagate through the system.
System architecture
At the centre were three platforms:
- GoHighLevel
- State, pipeline, calendars, SMS/email and workflow logic.
- Make
- Integration and orchestration between services.
- Vapi
- Live AI voice conversations.
Around them sat the acquisition sources, telephony, models, speech services, knowledge layer, newsletter infrastructure, website and reporting systems. The production stack ultimately touched 16 platforms.
- Rules deterministic logic
- AI a language model
- Human a person
- Lead sourcesMarketplaces · Meta · Website · Phone · Social
- Intake Rules
- Extraction from messy sources AI
- Normalisation and dedupe Rules
- GoHighLevel: CRM and state Rules
-
SMS / email Rules
- Make Rules
- Vapi voice agent OpenAI · voice · speech-to-text AI
- Live call AI
- Availability and knowledge lookup AI
- Post-call result AI
Direct booking Rules - GoHighLevel: pipeline state Rules
- Appointment booked Rules
- Human closer Human
Ten ways into one system
The system received prospects through ten intake channels. These included high-volume marketplace enquiries, Meta lead forms, paid-social landing pages, the main website, inbound telephone calls, social DMs and several manually triaged partner sources.
The important architectural decision was that these did not become ten separate sales systems. They converged into one model.
One discriminator
A custom lead-source field encoded both the relevant brand and acquisition path. Downstream logic could therefore ask:
Where did this prospect come from?
and derive:
- which brand they belonged to;
- which voice agent should call;
- which calendar should be used;
- which nurture track they should enter;
- which cadence variant should run.
That avoided duplicating entire CRM pipelines for every source.
The system had originally used separate 19-stage pipelines for brands. These were eventually collapsed into one shared 25-stage pipeline, making a new brand primarily a data/routing change rather than another duplicated architecture.
Store variation as data, not duplicated structure.
A new brand should require a new value and a few routing branches, not a cloned CRM architecture.
Making messy intake usable
Not every lead arrived cleanly. Marketplace leads, for example, often arrived as emails rather than structured webhook payloads. A typical flow was:
- Marketplace enquiry
- RulesMail rule
- RulesMake webhook
- AILLM extraction
- RulesSource lookup
- RulesDeduplication
- CRM
The LLM extracted fields such as:
- name;
- email;
- phone;
- comments;
- source identifier;
- URL.
The source identifier was then looked up to determine which brand owned the enquiry. If the system could not identify the source, it did not guess. It stopped and alerted the team. Likewise, parsing failures stopped the scenario before bad information reached the CRM.
Missing phone numbers were routed into a separate state and nurtured through email rather than accidentally entering the dial queue.
This mattered because AI orchestration is only useful when the state beneath it is reliable.
The CRM was the state machine
The core pipeline eventually contained 25 stages across four broad zones.
Lead states, by who owns them
-
Intake
Automation- New, with phone
- New, without phone
New leads entered with or without a valid phone number. Automation owned this stage.
-
AI-worked
AI- Callback requested
- Interested but not booked
- Unsuccessful AI follow-up
- Meeting no-show
- Short-term no-answer
- Not interested
- Text-only
- Undecided
The voice and messaging system owned most of these states.
-
Booked
Rules- Meeting booked
- Reminders and closer alert
Once a meeting existed, the lead transitioned toward human ownership.
-
Post-meeting
Human- Human closer
- Sale or loss
After the closer met the prospect, AI outbound calling stopped. The human sales process took over.
The data contract
About a dozen CRM fields formed the interface between GoHighLevel, Make and Vapi. Important examples included:
- Lead source
- Brand and acquisition source.
- Call outcome
- The structured result returned after each call.
- Consecutive misses
- How many unanswered attempts had occurred in sequence.
- Contact status
- Including whether booking succeeded or failed.
- Dial state
- Whether Vapi accepted the outbound request.
- Call record
- Summary, transcript, call type, number and timezone.
This is where the project becomes more interesting technically.
The voice model was not the system. It was one component producing structured inputs into a larger state machine.
The AI voice layer
Over the life of the system, 47 Vapi assistants existed across production, testing, previous versions and future iterations. Only around ten were typically live simultaneously.
47 assistants created~10 live at any one time
The system evolved through specialised assistants for different contexts. These included:
- First contact, per brand
- Paid-social first call
- Paid-social follow-up
- Callbacks
- Interested but not booked
- General follow-up
- No-show rebooking
- Missed-call callbacks
- Reactivation
- Inbound calls
- Newsletter-sourced leads
The reason for specialisation was simple:
Someone asking for a callback should not hear the same opener as a cold lead.
The system’s state gave the agent contextual knowledge before the call began.
Prompt architecture
The voice agents used long system prompts, typically around 5,000–16,000 characters. The structure became reasonably standardised:
- Role and persona
- Delivery rules
- Lead context
- Conversation flow
- Scheduling
- FAQ and objections
- Guardrails
The later scripts moved away from rigid pitching toward a more discovery-led interaction: understanding the prospect’s situation, explaining the model, reframing objections and then moving toward scheduling. Some prompts contained roughly 20 scripted objections.
We also spent considerable time on something much less impressive in a demo but critical in production: how the agent sounded.
Pacing, interruption handling, response latency, speech models, filler language and turn-taking all materially changed whether a conversation felt usable.
At the time, voice models were moving extremely quickly. Configurations that had taken considerable effort to make acceptable could become obsolete months later as latency and model quality improved.
What one lead actually experienced
One of the easiest ways to understand the architecture is to follow a marketplace enquiry through it.
-
0 sec
Enquiry submitted
A prospect submits an enquiry through a marketplace.
-
+ seconds
Parsed and deduped AI
The notification email reaches a mailhook. An LLM extracts the contact data and source identifier. The system determines the correct brand and checks the CRM for an existing record.
-
+ seconds
CRM state set Rules
The CRM creates or updates the opportunity, stamps its source and:
- sends the relevant welcome SMS;
- enters it into the appropriate nurture track;
- enrols it into AI calling.
-
+ ~15 sec
Outbound call initiated Rules
After a deliberately small delay, the dialer passes the prospect to the correct voice agent using a local number. In some flows, we deliberately sent SMS first or introduced a short pause rather than creating an unnaturally instantaneous phone call.
-
During the call
Live qualification and availability AI
The agent:
- confirms the enquiry;
- explains the offer;
- qualifies;
- handles objections;
- identifies the prospect’s location/timezone;
- queries live calendar availability;
- agrees a meeting time.
-
+ ~40 sec after the call
Transcript processed, appointment written AI
The transcript processor:
- classifies the result;
- saves the call;
- extracts the agreed appointment;
- writes it into the calendar;
- moves the prospect into the appropriate pipeline state.
-
T−24 hr
Prospect and closer reminders Rules
The system sends prospect reminders and alerts the assigned closer. If the prospect no-shows, separate email, SMS and AI voice recovery processes can begin.
-
Meeting
A human takes over Human
The hardest voice problem: trust the tool, not the model
One recurring failure mode came from scheduling. The agent was supposed to check live calendar availability before agreeing to a meeting.
But if the availability tool failed, or if the conversation diverged in an unexpected way, the model could still behave conversationally and accept a proposed time. To the prospect, the call sounded successful. Operationally, it wasn’t.
The downstream workflow would later attempt to write the appointment and discover that the agreed time did not exist.
That created a useful design lesson:
A conversationally plausible answer is not the same thing as a valid business action.
The system eventually treated booking as a two-part process. During the call: check and agree. After the call: validate and write.
The model checked live availability while speaking with the lead. Once the call ended, the transcript was analysed again, the agreed time was extracted in the prospect’s timezone and the appointment was written into the correct calendar.
That kept slower CRM writes out of the live conversation and made failed appointment writes detectable rather than silently lost.
During the callCheck and agree
- Lead on the call
- Asks city and timezone AI
- Requests availability AI
- Make parses the date and time AI
- GoHighLevel free-slots API Rules
-
Available: agree the time AINot available: propose alternatives AI
- End call
After the callValidate and write
- Transcript analysis AI
- Extract the agreed timestamp AI
- Write the appointment Rules
-
Written: booked stage RulesFailed: human alert Human
Follow-up was a system, not a retry button
A lead who failed to book was not simply called again forever. Pipeline state determined the next strategy.
Different stages had different:
- scripts;
- retry windows;
- call caps;
- cooling-off periods;
- voice agents.
The main stages used 8–16 attempts across roughly 7–10 days, depending on the state and brand. A typical no-answer pattern might begin aggressively before backing off:
A typical no-answer pattern
- 2 min
- 2 min
- 2 hr
- next day
- backs off
The point was to capture high-intent demand while it was fresh without blindly hammering every lead indefinitely.
A state-driven loop
- Call AI
- Outcome AI
- Booked Stop AI, send reminders, hand to the closer
- Interested Dedicated follow-up state
- Callback Callback state
- Unclear AI follow-up state
- No answer Add to the miss count, retry
- Not interested Exit the active cadence
A universal exit process removed leads from calling sequences when they:
- booked;
- entered closing;
- won;
- no-showed into another workflow;
- explicitly declined;
- called the business themselves.
This prevented the particularly bad automation failure of calling someone with a sales pitch after they had already booked.
Number reputation became a system constraint
One of the more important lessons was that the limiting factor was not always AI capability. It was telecommunications infrastructure.
Long-term aggressive calling eventually caused carriers to flag outbound numbers as spam. A long-term no-answer campaign was therefore shut down, and voicemail drops that had been built were left disabled.
The architecture adapted through:
- local numbers;
- a pool of 65 outbound numbers;
- drip throttling;
- per-stage call limits;
- reduced long-tail calling;
- greater reliance on email reactivation.
This is a good example of why an AI system cannot be designed in isolation from the environment it operates in.
The human handoff was deliberate
The objective was never to remove humans from the entire sale. The product was high ticket and consultative. A human closer remained more valuable once a prospect had committed enough attention to book a meeting.
So the system’s boundary became:
- Initial contact
- Qualification
- Routine objections
- Follow-up
- Appointment setting
- No-show recovery
- Selected reactivation
- State
- Routing
- Timing
- Calendars
- Retries
- Exits
- Alerts
- Closing calls
- Complex edge cases
- Relationship-sensitive scenarios
- Support
- Failed bookings requiring intervention
- The sales process after the meeting
The setter layer gradually reduced from four people. Three poorly performing setters left the process, while one moved into a full-time closer role. Setters were still occasionally used for edge cases that the system could not appropriately handle.
That was the point of the architecture:
Move humans up the value chain rather than forcing them to perform repetitive work the system could handle more consistently.
Reliability mattered more than the demo
A demo usually shows:
Lead arrives → AI calls → meeting booked.
Production has to answer:
What happens when something fails at every step?
The system eventually included patterns such as:
Retries
CRM writes and lookups retried automatically rather than dropping the run on transient errors.
Fail closed on bad data
Unknown source identifiers or extraction errors stopped processing instead of guessing.
Booking alerts
If a prospect verbally agreed to a meeting but the calendar write failed, an internal alert was generated for manual recovery.
Test isolation
Known test records could be diverted before triggering real calls.
Central state
Middleware reported facts; the CRM made business-state decisions.
The production system deliberately used incomplete-execution queues and retry handlers so failed integrations could be replayed rather than silently lost.
The system measured itself
Every AI call generated more than a transcript.
Post-call analysis
- Call outcome
- Booked time
- Main objection
- Whether the prospect appeared to detect the AI
- Their reaction if so
- Where the agent performed poorly
- Call summary
- Success score
Analytics API, per call
- Duration, result and agent
- Objection and recording
- Total cost
- Speech-to-text cost
- LLM cost
- Voice cost
- Platform cost
- Carrier cost
The post-call analysis created a feedback loop for prompt iteration. A custom analytics API also received per-call information, including a full cost breakdown.
The purpose was not just reporting. It allowed us to ask:
- Which agent is expensive?
- Which script fails most often?
- What are prospects objecting to?
- Where does the AI break?
- What does a booked meeting cost?
The CRM timeline then gave the closer a consolidated history containing AI call summaries and transcripts alongside SMS and email activity.
The feedback loop
- Call AI
- Transcript AI
- Structured analysis AI
- Outcome
- Objection
- AI detection
- Weak point
- Summary
- CRM and analytics Rules
- Prompt or flow change Human
- Next version
It was never “set and forget”
This was probably the most important operational lesson. The system evolved constantly.
Voice models improved. Speech latency changed. Prompting approaches improved. Business offers changed. Brands were added. Marketing channels changed. Call reputation changed. The CRM evolved.
What made sense in one Make scenario six months earlier could often be significantly simplified when a platform later introduced native functionality.
The technical audit itself found a system with 83 live CRM workflows from 158 built, 22 active Make scenarios in an account containing 163, and around 10 active voice assistants among 47 total versions.
Production footprint
- CRM workflows
- 83live of 158 built
- Make scenarios
- 22active of 163
- Voice assistants
- ~10live of 47 versions
- Outbound numbers
- 65
- Integrated platforms
- 16
- Intake channels
- 10
Live or activeBuilt over the system’s life
Some of that is the inevitable archaeology of a long-running production system. But it also reinforces something I now consider fundamental:
Automation needs ownership.
A system this interconnected does not become “finished” because the first version works. It needs monitoring, maintenance, experimentation, documentation and periodic simplification.
What went wrong
-
We inherited an undocumented system
The previous developer left virtually no usable documentation. Naming was inconsistent and often ambiguous. Even spelling errors became part of the implicit contract between workflows.
This made troubleshooting unnecessarily slow and changes unnecessarily risky.
What changed
Architecture documentation, naming discipline and future maintainability became explicit requirements.
Later, giving AI coding tools direct access to APIs and the surrounding system also dramatically improved the ability to trace and understand dependencies.
-
The agent could sound right while being operationally wrong
If the availability tool failed, an LLM could still conversationally accept a time. The prospect believed they had booked. The system had not.
What changed
Booking was treated as a controlled transaction rather than trusting conversational output. Failures became explicit states with alerts.
-
Timezone handling caused bad calling windows
An early configuration could result in East Coast prospects receiving calls at inappropriate local hours.
What changed
Calling was standardised around a safer operating window, while agents captured timezone context for scheduling. The audit still identified local-time scheduling as an area for further improvement.
-
Carrier reputation limited theoretical scale
More calling was not automatically better. Long-term campaigns damaged number reputation and reduced deliverability.
What changed
Long-term calling was reduced, voicemail drops stayed disabled and more reactivation moved to email.
-
Old integrations accumulated debt
The technology landscape moved quickly. Some Make scenarios that had been sensible when built could later be replaced or dramatically reduced as platforms developed better APIs or native integrations.
What changed
Regular architectural review became part of maintaining the system rather than assuming existing workflows remained optimal.
-
Dependencies can fail invisibly
At one point, a knowledge-base backend was disabled and dozens of agents lost their fallback capability without producing an obvious system-wide error. The audit identified dependency health checks as a missing safeguard.
What changed, conceptually
A dependency being reachable needs to be monitored separately from whether the overall workflow appears “on”.
Results
The business was collecting roughly $200,000 per month before the system matured. As marketing increased lead volume, the AI-enabled sales system became the infrastructure that allowed the business to handle that growth without scaling the setter team alongside it.
At maturity, the system supported:
- AI calls in a peak month
- ~50K
- meetings booked autonomously each month
- ~400
- sustained cash collection across the sales operation
- ~$1M/mo
The system did not create that revenue in isolation. It supported the scale required to handle a much larger volume of inbound demand consistently.
Monthly cash collection
- Before the system matured
- ~$200K
Acquisition and automation channels added progressively
- Mature operating period
- ~$1M
Every lead could now move through a defined process: near-immediate first contact, persistent follow-up, structured CRM state, automated booking and a clear handoff to a human closer.
The result was not simply more calling capacity. It was a sales infrastructure that could absorb substantially more demand without allowing lead management to become the bottleneck.
What I learned
Three lessons have carried into almost everything I have built since.
Design for the system you will have to maintain
The quickest implementation is not necessarily the quickest system. Ambiguous names, undocumented dependencies and duplicated logic create a tax on every future change.
The more capable AI coding becomes, the more tempting it is to create quickly. That makes architecture and documentation more important, not less.
Separate intelligence from authority
An LLM can be excellent at interpretation without being the right mechanism for making every decision.
“Thursday sometime around two.”
Whether two o’clock is actually available and what state transition follows.
Giving AI the parts it is good at while constraining authority elsewhere made the system much more reliable.
Automation is an operation, not an installation
This system required:
- Monitoring
- Prompt iteration
- Testing
- Carrier management
- Workflow maintenance
- Reporting
- Cleanup
- Documentation
- Model upgrades
- Process changes
The idea that a business can “turn on AI” and leave it untouched misunderstands the nature of production automation. A system that touches customers and revenue needs an owner.
What I would build differently today
If I were rebuilding the system from zero today, I would simplify it substantially.
Reduce platform dependence
Several functions that historically required Make or complicated CRM workflows could now be implemented more cleanly in custom services or newer native integrations.
Make contracts explicit
Instead of state depending on free-text fields and exact strings, I would use enumerated states and stable identifiers wherever possible.
Design local-time calling from the beginning
Each lead’s timezone would become part of scheduling and cadence immediately rather than being captured but only partially operationalised.
Add dependency health monitoring
Every external tool would have explicit health monitoring: voice, calendar, knowledge base, CRM and model endpoint.
Treat observability as a first-class system
Errors, costs, latency, conversions, call quality and dependency failures would live in one consolidated observability layer rather than being distributed across tools.
Prune continuously
Production systems accumulate old versions. Retirement should be part of the development lifecycle rather than a housekeeping exercise undertaken later.
Full technology stack 14 tools
| Layer | Technology | Purpose |
|---|---|---|
| CRM / state | GoHighLevel | Pipeline, workflows, calendars, SMS/email |
| Integration | Make | Data movement, utility AI calls, API orchestration |
| Voice | Vapi | Outbound/inbound AI calling |
| Telephony | Twilio | Numbers, carrier layer, SMS |
| LLM | OpenAI | Conversation, extraction, date/time interpretation |
| Voice | ElevenLabs | Text-to-speech |
| Transcription | Deepgram | Speech-to-text |
| Knowledge | Voiceflow / Claude | Off-script knowledge lookup |
| Acquisition | Meta | Lead forms, paid social, conversion signals |
| Web | WordPress / Elementor | Landing pages and booking |
| Social | ManyChat | DM/comment automation |
| Nurture | Beehiiv | Long-term email nurture |
| Reporting | Google Cloud Run | Call analytics and cost ingestion |
| Operations | ClickUp | Downstream delivery operations |
The phone call was the visible part.
The interesting engineering happened around it. A production AI sales system needed to know:
- who the prospect was;
- why they had entered;
- which brand they belonged to;
- what had already happened;
- which conversation was appropriate now;
- when to call again;
- when to stop;
- whether the calendar was genuinely available;
- whether the booking actually succeeded;
- when to bring in a person;
- and what to do when any dependency failed.
That is the difference between an AI demo and an operating system.
The model handled the conversation. The surrounding architecture made it useful to the business.