Projects

Case study

Link acquisition operating system

Rebuilding a six-person manual operation around software, rules and AI

2.5% → 28.5% of orders live within 30 days · 25,883 domains assessed · 87% of quality decisions made without a person

A link-acquisition business had delivered thousands of backlinks for agencies and direct clients over several years. Clients ordered links to improve the authority and search visibility of their websites; the operation sourced suitable publishers, vetted their quality, negotiated placement rates and managed each order through publication and payment.

The underlying business was healthy. The operating model was not.

Prospecting, vetting, outreach, negotiation and fulfilment were spread across SOPs, Airtable, Pitchbox and shared inboxes, supported by a six-person team. By the time I started the rebuild, average orders were taking roughly 80–88 days to go live and only 2.5% were being delivered within 30 days.

The team itself consisted of 1 manager, 1 outreach operator and 4 fulfilment staff.

The objective was not simply to “add AI”. It was to redesign the operation so it could:

  • source new publishers at scale;
  • enforce a consistent quality bar;
  • negotiate rates systematically;
  • fulfil orders faster;
  • lower tooling and staffing costs;
  • preserve human judgement where mistakes were expensive;
  • and create a system that could be measured and improved continuously.

I handled the project end to end: process discovery, system specification, architecture, data model, cost modelling, implementation, testing, deployment and ongoing iteration.

At a glance

domains automatically assessed
25,883
of quality decisions made without a person
87%
orders live within 30 days
2.5% →28.5%
links live in the first four months
915

The platform also accumulated 1,511 won publishers with agreed rates, stored 15,564 rate cards, imported 71,080 historical emails and monitored more than 16,000 live links.

Where we started

Before the new platform, the operation ran through:

  • Two written SOPs
  • Airtable
  • Pitchbox
  • Four shared mailboxes
  • Manual publisher vetting
  • Manual negotiation
  • Manual placement selection
  • Manual payment tracking

The underlying business was profitable and established. Since 2022 it had delivered more than 16,000 links for 42 clients at healthy margins. The process around it was the problem.

Prospecting was manual

A team member worked through lists in Pitchbox, removed obvious bad-fit sites, found contact details and manually checked each domain against the agency’s quality criteria. That included factors such as:

  • language;
  • recency of publishing;
  • article volume;
  • whether the site openly sold links;
  • PBN indicators;
  • niche restrictions;
  • basic quality thresholds.

Negotiation was deterministic on paper, manual in practice

The SOP already contained:

  • DR-based price bands;
  • a three-step offer ladder;
  • fixed follow-up timing.

Humans were manually executing rules that were largely predetermined.

Fulfilment lived in a master sheet

Once a client ordered a link, someone manually selected a publisher, emailed them, chased the placement, tracked the live URL and handled payment. As volume increased, that process became harder to manage consistently.

The clearest symptom was delivery speed

The decline was visible quarter by quarter.

Orders live within 30 days, by order quarter

By April and May 2026, average orders were taking 80–88 days to go live, and none went live within 30 days.

Defining the system

The goal

The brief was straightforward in business terms:

Pitch thousands of sites per month with minimal human involvement, preserve the quality bar, and make the publisher database something the system could act on directly.

The deeper challenge was that this required redesigning the process, not automating individual tasks. The existing operation had grown around people, tools and habits. The new one needed explicit rules.

One week before writing code

I spent roughly a week specifying the system before implementation began. That included:

  • reviewing the existing SOPs;
  • interviewing the manager who ran the process;
  • mapping the actual workflow rather than assuming the SOP matched reality;
  • defining every major decision point;
  • designing the database schema;
  • identifying external data sources;
  • modelling API limits and likely costs;
  • estimating system capacity;
  • separating deterministic decisions from judgement calls;
  • defining human approval points.

The design specification was signed off before implementation, followed by an engineering plan that fixed the starting schema, module boundaries and thresholds.

That sequence became one of the most important lessons from the project:

Process discovery is part of the build.

Skipping it would have produced a faster first version of the wrong system.

Six decisions made before the build

Several policy decisions were resolved before they could become accidental software behaviour.

  • Niche fit would be hybrid

    Sites would be broadly classified during vetting, but suitability for a specific client would be decided later during order matching. A good site should not be rejected globally because it was unsuitable for one client.

  • PBN detection would be balanced

    Obvious networks could be rejected automatically. Borderline cases would go to human review rather than being silently discarded.

  • Traffic concentration mattered

    A site could have headline traffic and still be low quality if one page carried almost all of it.

  • Traffic decline needed context

    A decline in clicks alone was not enough. Traffic loss combined with ranking deterioration was more meaningful than traffic movement in isolation.

  • Very high DR would not automatically reject a site

    DR 90+ could be deprioritised without being excluded.

  • Re-pitching needed a cooldown

    Publishers should not be repeatedly contacted without enough time between campaigns.

These decisions were part of the original system specification.

The architecture

The platform was built as a custom Python application around a MySQL database. The core stack included:

  • Python
  • MySQL
  • FastAPI
  • HTMX
  • APScheduler
  • Ahrefs
  • DataForSEO
  • Hunter
  • Instantly
  • OpenRouter / Claude
  • Legacy IMAP/SMTP mailboxes

But the more important architectural decision was not the technology stack. It was the integration model.

The database was the integration point

Every worker interacted through the database. A worker would:

  1. Claim an item in a known state Rules
  2. Perform one bounded piece of work Rules
  3. Write the result Data
  4. Advance the state Data

Nothing needed to directly call every other subsystem. That gave the system several useful properties:

  • restarts were safe;
  • work could resume from the last committed state;
  • jobs were inspectable;
  • failures could be retried;
  • workers could evolve independently;
  • every decision had a persisted history.

The design principle was:

State moves through the database. Workers do not depend on each other being alive at the same moment.

That made the platform much easier to operate than a long synchronous automation chain.

  • Rules deterministic logic
  • AI a language model
  • Human a person
  • Data stored or bought data
  1. Sourcing
  2. SERPs Data
    Competitors Data
    Rate cards Data
  3. Normalisation Rules
  4. Deduplication Rules
  5. Vetting engine Rules
  6. Rules Rules
    AI read AI
    Data APIs Data
  7. Decision ledger Data
  8. Verified contact Rules
  9. Outreach Rules
  10. Reply classifier AI
  11. Negotiation engine Rules
  12. Won publisher database Data
  13. Orders
  14. Matching engine Rules
  15. Human approval Human
  16. Placement Rules
  17. Live
  18. Payment Human
  19. Monitoring Rules

Rules decide. AI interprets.

This was the most important separation in the platform. The system used AI, but it deliberately did not use AI for every decision.

Deterministic logic handled
  • Pass/fail thresholds
  • Prices
  • Maximum allowable costs
  • Traffic requirements
  • DR requirements
  • Niche matching rules
  • Cooldowns
  • Offer ladders
  • Margin constraints
  • Workflow state
  • Order matching
AI handled
  • Reading and classifying content
  • Deciding whether a site looked like an editorial publisher
  • Extracting structured meaning from publisher replies
  • Identifying intent
  • Rewriting or personalising language
  • Parsing human communication
Humans handled
  • Ambiguous review cases
  • Large or unusual commercial decisions
  • Publisher payments
  • Final placement approval
  • Exceptions
Use intelligence where interpretation is required. Use rules where certainty is available.

The rule was explicit:

The LLM reads and writes words. It does not set prices or invent business facts.

The vetting pipeline

Every sourced domain could pass through up to eight stages. The key was that they were deliberately ordered by cost, not by perceived importance. The early checks removed obvious failures cheaply. Expensive data was only purchased for the smaller group that survived.

The broad sequence was:

  1. RulesBlocklists
  2. AIEditorial / brand-safety check
  3. DataMarket traffic
  4. DataDomain Rating
  5. AIContent quality
  6. DataPBN / advanced quality analysis
  7. RulesVerified contact
  8. Outreach

Of roughly 25,000 domains entering the pipeline, only a small fraction reached cold outreach.

Why ordering mattered

Running every check on every domain would have been wasteful. For example:

  • DR could be checked cheaply and reject a large share early;
  • more expensive graph and traffic-history analysis should only run on survivors;
  • contact enrichment was pointless for a site that already failed quality;
  • outreach should only happen after qualification.

The system therefore followed a simple principle:

Spend progressively more only as confidence increases.

Domains remaining after each stage

  1. Domains in24,990
  2. RulesFree filters23,072
  3. AICheap AI read12,707
  4. DataTraffic4,913
  5. DataDR4,300
  6. AIContent3,165
  7. DataExpensive quality checks878
  8. RulesContact, then outreach659
Cost per candidate rises as candidate volume falls.

That architecture reduced expected API spend substantially. The original design estimate was roughly $1,700–$2,200 per month. Actual metered API usage since launch totalled only $891 across the first four months.

Every rejection had to explain itself

A pass/fail result alone was not enough. The platform recorded:

  • which stage made the decision;
  • which check triggered it;
  • the actual measured value;
  • the threshold at the time;
  • the data source;
  • whether a human overrode it.

That produced a decision ledger. For example:

Decision ledger entry

Traffic
640
Threshold
1,000
Decision
Reject
Source
Ahrefs
Human override
No

This became extremely valuable later. Because the historical measured value was preserved, a threshold could be changed and tested against previous data without paying to re-run the original check.

The thresholds were not sacred

The first values came from the SOPs. Production data changed some of them. For example, the initial minimum traffic threshold was 1,000. Later, the system distinguished between:

  • vetting threshold;
  • guest-post placement threshold;
  • link-insert placement threshold;
  • warning ranges.

Likewise, the original trend-detection logic did not perform well enough in practice. It could reject healthy niche sites while missing actual deterioration. That logic was replaced with a better signal using recent organic-traffic history against earlier periods.

That is an important theme throughout the project:

The specification created the starting model. Production evidence refined it.

Cost engineering

One of the more interesting parts of the system was that API cost became a measurable engineering constraint. Every external request recorded:

  • Service
  • Endpoint
  • Units consumed
  • Dollar cost
  • Domain
  • Timestamp

The platform accumulated hundreds of thousands of metered call records. That made it possible to identify expensive logic rather than guessing.

Batch everything that can be batched

Ahrefs was a good example. A one-domain call carried a large fixed request overhead. Batching many domains into one request dramatically reduced units per domain. The system therefore:

  • batched DR requests;
  • batched traffic lookups;
  • used free endpoints where possible;
  • deferred expensive graph checks until late;
  • avoided rechecking values already known.

The cheapest model was not actually the cheapest

The original system used a mixture of Claude and a self-hosted Qwen model on rented GPUs. On paper, the self-hosted model looked economical. In production it created operational problems:

  • GPU capacity was inconsistent;
  • cold starts could take 6–20 minutes;
  • queues could grow faster than they drained;
  • one stalled queue faced a projected 37-hour processing time.

The system was eventually moved to OpenRouter, with the model selected per task and Claude retained as fallback. The result was dramatic. LLM cost fell from roughly $89 in June to $0.98 in September, despite September processing more than 10,000 LLM calls.

LLM cost per month

  1. $89.47June
  2. $49.33July
  3. $9.06August
  4. $0.9810,559 callsSeptember

Getting the data was harder than building the screens

Four years of operating history existed across multiple systems:

  • Airtable
  • Excel exports
  • Pitchbox
  • Four mailboxes
  • Payment records

The challenge was not simply importing it. The data had to become trustworthy enough for software to make decisions from it.

Reconstructing historical truth

The import exposed several problems.

  • Orders

    The first Airtable import collapsed thousands of orders into 101 records because the wrong identifier had been treated as unique. That had to be rebuilt around the actual record ID.

  • Publisher rates

    The real negotiated publisher rates were not stored cleanly in the publisher table. They had to be reconstructed from historical orders and what had actually been paid.

  • Missing DR

    Hundreds of publishers had no usable DR on file. Those values were backfilled from live data.

  • Conflicting exports

    An Excel publisher export was not actually a superset of Airtable. Missing values could be added from it, but replacing existing records wholesale would have destroyed valid data.

  • Historical inboxes

    More than 71,000 emails dating back years were pulled into one searchable history and associated with publishers by contact, domain, subject and thread.

This part of the project reinforced another lesson:

Automation magnifies bad data just as effectively as good data.

Before the system could automate decisions, the underlying history had to be repaired.

Sourcing new publishers

One of the biggest business constraints before the build was that new publisher acquisition had almost stopped. The new platform added multiple sourcing methods. These included:

  • Reseller rate cards

    Thousands of publisher domains could be imported from reseller inventory and passed through the same quality pipeline.

  • Competitor expansion

    The platform found sites competing for similar search visibility to known good publishers.

  • Publisher catalogues

    Bulk partner datasets could be ingested and vetted.

  • SERP harvesting

    Search results for niche queries produced candidate publishers at very low cost.

  • Guest-post lists

    Existing lists could be normalised and passed through the standard pipeline.

  • Backlink mining

    This was also tested, though much of the resulting inventory was poor quality.

The platform ultimately sourced tens of thousands of domains across these channels.

domains sourced
27,910
duplicates skipped before paid checks
2,200+

Deduplication came before spend

A domain already in the database should never re-enter paid vetting. That sounds obvious. At scale, it matters.

The system skipped more than 2,200 duplicates before paying for additional checks. That is another small example of why economic design had to happen at the architecture level rather than after the fact.

Outreach

Only domains that passed the quality bar reached cold outreach. This reversed the previous operating model, where sites could be contacted before they had been fully qualified. That change saved sending capacity and avoided wasting attention on publishers the business could never use.

Outreach ran through warmed mailboxes in Instantly, with multiple opening variants tested in production. The three main variants produced reply rates in the low-to-mid 20% range, each outperforming the previous manually operated process, which had typically produced roughly 8–15% reply rates.

AI read the reply. Rules decided what happened next.

Replies were classified by an LLM into a fixed set of intents. Examples included:

  • Gave a price
  • Interested
  • Not interested
  • Out of office
  • Redirected to another contact
  • Accepted the offer
  • Guest posts only
  • Link inserts only

The LLM also extracted information such as the quoted price, currency and VAT. But it did not decide how much the company should pay. That was deterministic.

This separation became one of the cleanest examples of the overall architecture:

Use AI to understand the message. Use code to decide the business action.

Negotiation was the SOP turned into software

The existing negotiation SOP already had:

  • maximum prices by DR band;
  • separate ceilings for guest posts and inserts;
  • three offer steps;
  • follow-up timing.

The platform encoded those rules directly. A simplified negotiation flow looked like:

  1. Publisher quotes
  2. Price read from the reply AI
  3. Price within target? Rules
  4. Yes: win Rules
    1. No: offer step 1 Rules
    2. Step 2 Rules
    3. Step 3 Rules
    4. Firm final price? Rules
    5. Yes: accept Rules
      No: continue or human review Human

Extreme outlier quotes were escalated rather than blindly following the ladder. For example, countering a $2,000 ask with $250 repeatedly made the automation look unserious. Quotes far beyond the expected band were therefore routed to a manager.

The negotiation results

of 87 publishers won in automated negotiation
41
lower guest-post pricing against opening quotes
~25%
lower link-insert pricing against opening quotes
~20%

Of 87 publishers that entered automated negotiation, 41 were won. Across publishers with recorded opening quotes, guest-post pricing fell by approximately 25% and link-insert pricing by approximately 20%.

A won publisher was not treated as a one-off transaction. Its agreed price became a reusable rate card. That meant the real asset was not the current placement. It was a growing database of publishers with known commercial terms.

Fulfilment

Once a client order arrived, the system had to select the right publisher. This required both hard exclusions and ranking preferences.

Hard rules: a publisher is excluded if it is

  • Already linking to the target client
  • Failing the required margin
  • Below the required DR
  • Below the required traffic
  • In the wrong market
  • Previously used for the same target
  • No longer active

Ranking: survivors are ordered by

  • Niche relevance
  • Cost

The platform generated three ranked suggestions for each order. Managers could override the recommendation, but they rarely needed to. The matching engine selected approximately 93% of placements without override.

Monitoring did not stop when the link went live

Since September, published links were rechecked automatically. The system differentiated between:

  • Live
  • Dropped
  • Blocked / cannot verify
  • Not yet checked

This distinction mattered. A site hidden behind a bot challenge should not automatically be classified as a dropped link. The system required plausible evidence before marking a placement as gone.

That reflects a broader theme of the architecture:

Unknown is not the same thing as failed.

Client portal

The two largest partners were given access to a read-only client portal.

They could

  • View their own orders
  • See live links
  • Filter results
  • Export data

They did not see

  • Internal publisher proposals
  • Supplier costs

This gave the system another role: not just operational infrastructure, but client-facing reporting.

The platform also became the historical record

One of the less visible but important outcomes was that operational history stopped being fragmented. The system consolidated:

  • Orders
  • Publishers
  • Rate cards
  • Email threads
  • Payments
  • Placement history
  • Quality measurements
  • Link status
  • Client delivery

The historical inbox alone contained more than 71,000 messages. This meant decisions could be made from the actual history of the publisher relationship rather than relying on institutional memory.

Human gates remained where money moved

The platform was designed to automate legwork, not remove accountability. Two major human approval points remained:

Placement approval

A person approved the selected publisher before commitment.

Financial approval

A person approved payment or prepayment decisions.

An order only closed when the link was live and paid, in either order.

Human-in-the-loop should be based on the cost of being wrong, not on whether automation is technically possible.

From six people to two

6 people

1 manager · 1 outreach operator · 4 fulfilment staff

  • SOPs
  • Airtable
  • Pitchbox
  • 4 mailboxes
  • Manual vetting
  • Manual negotiation
  • Manual fulfilment

2.5%live within 30 days

1 full-time operator, 1 part-time manager

Custom platform

System handles

  • Sourcing
  • Vetting
  • Outreach
  • Reply interpretation
  • Negotiation
  • Matching
  • Follow-up
  • Link monitoring
  • Historical data lookup

Humans handle

  • Approvals
  • Payments
  • Edge cases

28.5%live within 30 days

The manager remains involved primarily around payments, approvals and exceptions. The operator manages the flow rather than manually performing every task.

The reduction in staffing was not the only objective, but it was a meaningful consequence of redesigning the operation.

What went wrong

The platform improved because real production usage exposed assumptions the original specification could not.

Initial approachWhat happenedWhat changed
First trend modelHealthy sites rejected, real decline missedRebuilt around historical traffic
Self-hosted LLMCapacity delays and long queuesOpenRouter and task-specific models
Bot-like requestsPublisher security blocked checksChrome-like fetching
Highest historical metricDeclining sites looked healthyLatest valid measurement
Rigid negotiationAbsurd counters on extreme quotesHuman escalation threshold
Loose email threadingReplies lost or duplicatedExplicit thread and reply rules
  1. The first traffic-trend logic was wrong

    The original method could flag healthy sites while missing actual deterioration.

    What changed

    The signal was replaced with a more robust traffic-history comparison. It became both cheaper and more accurate.

  2. The crawler looked too much like a bot

    Publisher security plugins blocked a large number of assessments.

    What changed

    Requests were modified to behave more like a normal Chrome browser, eventually including Chrome-like TLS behaviour. Read success improved materially.

  3. Self-hosted AI was operationally expensive

    The model itself looked cheap. The unavailable GPU capacity, cold starts and queues were not.

    What changed

    LLM traffic moved to OpenRouter with model selection by task. Per-call cost fell dramatically and operational reliability improved.

  4. Historical rates were not where we thought they were

    The publisher table did not contain the true negotiated rate history.

    What changed

    Rates were reconstructed from actual historical order payments.

  5. Reading the highest metric hid decline

    Using the maximum historical traffic or DR made deteriorating publishers look healthy.

    What changed

    The system now reads the latest valid measurement instead.

  6. Estimated traffic let weak publishers through

    Estimated values were not reliable enough for placement decisions.

    What changed

    The system moved to measured traffic and treated “unmeasured” as unresolved rather than passing it.

  7. Deterministic negotiation still had edge cases

    A rigid offer ladder could behave badly when the opening price was extreme or explicitly non-negotiable.

    What changed

    • Extreme quotes escalate.
    • Firm prices can be accepted after the ladder.
    • The system never counters above the publisher’s own asking price.
  8. Email threading was harder than it looked

    Real mailboxes contained:

    • Redirects
    • Personal-address replies
    • Old threads
    • Out-of-office messages
    • Attachments
    • Duplicate sends

    What changed

    Threading and reply classification became more explicit, with stored classifications, redirect workflows and duplicate-send protection.

How iteration actually worked

The production loop became simple:

  1. A team member encountered a real problem.
  2. The issue was traced against production data.
  3. The root cause was fixed.
  4. A regression test was added.
  5. Any affected historical data was repaired safely.
  6. Business-policy changes were separated from engineering fixes.

That last point matters. Not every production issue should immediately become code. Sometimes the system is working correctly and the business rule itself needs to change.

375 releases in four months

The platform shipped 375 releases in its first four months. The first two days produced the foundations, vetting, outreach, negotiation, fulfilment, console and scheduler. The rest came from production use.

The test suite grew from 112 tests at the end of day one to 272 after day two and 1,696 by the October audit.

  1. Day 1

    Foundations and vetting

    112tests
  2. Day 2

    Outreach, negotiation and fulfilment

    272tests
  3. Month 2

    Production workflows expand

  4. Month 3

    Portal, monitoring and data repairs

  5. Month 4

    375 releases

    1,696tests
The first version created the system. Production use made it good.

Results

The platform reversed a delivery decline that had been worsening for more than a year.

In the quarter before launch, only 2.5% of orders went live within 30 days. In the first full quarter after launch, that rose to 28.5%, more than an 11× improvement.

Orders live within 30 days, by order quarter

At the same time:

domains automatically assessed
25,883
quality decisions recorded
163,566
of them made without a person
87%
links live in the first four months
915
of the inherited backlog delivered
75%
of placements selected without manager override
93%

The operating model also changed from a six-person team to one full-time operator and a part-time manager, with software performing most of the repetitive legwork.

Links going live per month, 2026

  1. 109Apr
  2. 113May
  3. 212Jun
  4. 358Jul
  5. 231Aug
  6. 260Sep

The result was not simply automation. It was a more measurable, lower-overhead operation that could acquire new supply, process more opportunities and deliver substantially faster than the manual process it replaced.

Cost and tooling

The previous stack

  • Pitchbox: roughly $600/month
  • Airtable: roughly $25/month
  • Six-person team at approximately $1,500/month per person

The new dedicated recurring stack

  • Instantly: $97/month
  • Sending inboxes: ~$50/month
  • Server and hosting: ~$20/month
  • One sending domain: negligible annualised cost
  • Hunter usage
  • LLM usage, down to roughly $1/month
  • Metered data and API costs as required

Tools shared with other work, such as Ahrefs and Claude Code, are left out of both columns, so this is not a net-savings figure.

The platform materially reduced both staffing and dedicated software overhead while increasing throughput.

What I learned

Understand the process before automating it

The one-week specification phase probably saved more time than it cost. The existing SOPs were necessary but insufficient. The important work was understanding:

  • where judgement actually existed;
  • where people were merely executing rules;
  • where data was unreliable;
  • where money changed hands;
  • where mistakes would be expensive.

Only then could the automation boundary be designed intelligently.

Economics belong inside the architecture

API spend was not a line item to optimise later. It changed:

  • Check ordering
  • Batching
  • Provider selection
  • Caching
  • Deduplication
  • Model choice
  • When expensive data should be purchased

That is why the system could assess tens of thousands of domains without approaching the original cost estimate.

Production feedback is part of the build

The first version was not a failure because later logic changed. The later logic was better because the first version generated evidence. The system improved through:

  • False positives
  • Blocked requests
  • Strange publisher replies
  • Pricing edge cases
  • Bad historical records
  • Real operator behaviour

Each issue became a rule, test or architectural change. That is why the final platform looked different from the original specification while still preserving its core principles.

What I would build differently today

  • Build observability even earlier

    The decision ledger and API metering became extremely valuable. I would make those first-class from day one.

  • Keep human review reasons structured

    Any human override should require a fixed reason code plus optional commentary. That makes human judgement measurable too.

  • Separate policy from code more aggressively

    Thresholds, commercial rules and offer ladders should increasingly live in configuration rather than requiring releases.

  • Design email threading as its own subsystem

    Email looks simple until it becomes operational infrastructure. Thread identity, redirects, bounces, out-of-office states, attachments and duplicate protection deserve explicit modelling.

  • Make sourcing continuous

    The next logical step is turning sourcing from a manually initiated activity into an always-on supply engine driven by saved seeds and capacity.

Technology stack

LayerTechnology
ApplicationPython
DatabaseMySQL
APIFastAPI
UIHTMX
SchedulingAPScheduler
SEO / link dataAhrefs
Search / traffic dataDataForSEO
Contact discoveryHunter
OutreachInstantly
LLM routingOpenRouter
ModelsClaude and other task-specific models
Email historyIMAP / SMTP
DevelopmentClaude Code

The automation was not the point.

The important change was turning an implicit human operation into an explicit system.

Before the build, the business logic lived in

  • SOPs
  • Spreadsheets
  • Individual judgement
  • Inbox history
  • Habits
  • Memory

After the build, those decisions became

  • States
  • Thresholds
  • Recorded measurements
  • Approval gates
  • Deterministic rules
  • Structured exceptions
  • AI

    AI was useful where language and ambiguity existed.

  • Rules

    Code was better where the rule was known.

  • Human

    Humans remained where context, money or judgement made the cost of an incorrect decision meaningful.

That combination was what made the operation scalable.

The biggest gain did not come from replacing people with AI. It came from understanding the process well enough to decide what should be software, what should be AI and what should still require a person.