Visitor intelligence research

How Should an AI Agent Safely Find Email? A Founder's $47,000 Lesson on okki go Workflow

2026-09-18 · Erin Watanabe

Tuesday, 6:47 AM, February 2024

I wasn't dressed yet. The coffee had gone cold. Postmaster Tools was showing that our main sending domain had slipped from "High" to "Medium" reputation somewhere between 2 AM and 6 AM. That's roughly 90% of our outbound landing in spam for the rest of the day, I thought. Then the real thought hit: I'd onboarded two new SDRs three weeks earlier, and both of them had quota calls on my calendar for that afternoon.

This was the second time I'd personally torched an outbound channel. The first one cost me around $47,000 over three years. This one was about to cost me a lot more, because we were still running the same playbook.

How we got there (2021–2023)

When I took over outbound in 2021, "the team" was me, one salesperson, a LinkedIn Sales Navigator seat, and a 4,000-row Google Sheet. By Q2 we were buying records from an enrichment provider — a two-thousand-lead contract, priced per verified contact.

The logic felt obvious: more emails in the pipe → more replies → more pipeline. So we scraped, bought, enriched, and pushed everything into a sequencer.

Three questions I never asked, that I ask by default now:

  • Are these addresses verified at the point of purchase, or after the fact?
  • Am I paying the same per-record price for clean data and for garbage?
  • Where did the data come from — is it opt-in, publicly sourced, or scraped from a 2019 Excel file that changed hands eight times?

I asked none of these. Between 2021 and 2023, across enrichment API calls, duplicates, and dead records I kept buying because the CRM charged per seat, I burned roughly $47,000 on leads we never should've uploaded in the first place. Give or take. It was $43K according to the first invoice run — no, $47K by the end-of-quarter reconciliation. I've stopped trying to remember the exact number.

The mistake wasn't buying data. It was trusting the sequence.

Here's the part that hurt most when I finally audited it: we were verifying. We ran every list through an SMTP check and dropped the obvious hard bounces.

But the order was wrong. We bought first, then verified. That means:

  • We paid full price for records that failed verification.
  • We had no signal at purchase time about whether a source was good or rotten.
  • We couldn't trace a bad batch back to a provider because we never logged where each record originated.

When we finally pulled the numbers on a Q3 2022 audit, 22% of the addresses we'd bought were hard bounces, and another 12% were catch-alls that we had no business emailing cold.

Everything I'd read about B2B prospecting said "buy bigger lists, then clean them." In practice, the opposite worked: buy smaller, verify inside the enrichment call, and clean as you go.

I wrote the first version of our checklist in Notion on a Friday night. It was six bullet points and it still lives in our onboarding doc.

February 2024: Google rewrote the rules

If you were running cold email at any volume that February, you remember the bulk-sender requirements. Google made three things non-negotiable for anyone sending roughly 5,000+ messages a day to Gmail addresses:

  • Spam complaint rates had to stay under 0.30% — and yes, that threshold is real and documented in Google's Postmaster guidelines.
  • One-click unsubscribe in the message itself.
  • DMARC alignment on your sending domain.

We were under 5,000/day, technically. But we were on a shared sending IP with a vendor, and their other customers were not. So we got swept up anyway.

Our first batch after the change hit a 0.4% complaint rate in 72 hours. That's not catastrophic in isolation. But Google's machine doesn't look at each batch in isolation — it looks at your domain's history. Our reputation dropped and it took six weeks to recover.

I'm not going to pretend I understood all of this in February 2024. I didn't. A friend who runs deliverability for a competitor walked me through it over two calls. That's another thing I'd tell any founder: you don't need to be a deliverability expert, but you need someone in your orbit who is.

Then the AI personalization embarrassed us

We'd started dabbling with AI-written first lines in late 2023. Most of it worked. Some of it did not.

One incident stands out. Across about 50 outbound emails, the model wrote variations of:

"Saw you expanded the team from 8 to 12 in Q3 2022 and opened an EMEA office..."

None of that was true. We had no source for headcount. There was no EMEA office. And half the people we were emailing had left the company the previous year, because we were pulling from a stale enrichment API.

A prospect replied with a subject line that just said: "Where did you find this?" I still think about that email. Not because it was rude — it wasn't. Because it was a reasonable question and I didn't have a good answer.

That's the core of how an AI agent should safely find email. The agent can absolutely find a valid address. The failure mode isn't the lookup. It's the confident assertion that comes after the lookup, when the model has no idea whether the surrounding facts are true.

What we rebuilt around okki go

I want to be careful here, because the rebuild wasn't a tooling decision. It was a priorities decision. The tooling followed.

The four things we decided we would not compromise on:

  1. Find, then enrich, then verify — in a single pass. Not three tools stacked, not a batch that gets cleaned at the end of the week.
  2. Waterfall everything. If one provider doesn't have a working address, try the next — inside the same workflow, with verification at each step.
  3. A human stays in the loop on every batch. Not "approve" as a checkbox. A person who reads the personalized line and can say "no, that's wrong."
  4. Every record carries provenance. Source, date acquired, verification result. If we can't answer "where did this come from," we don't send.

We ended up running most of this through okki go (the product is okkigo, sometimes written okki-go — I've seen it spelled four different ways). What made it work for us wasn't any single feature. It was that the tool was designed around the assumption that a human reviews what the agent does, not the other way around. That's an unusual design choice in this category and it's the one I'd look for again if we had to switch.

The waterfall enrichment piece mattered too — a modest improvement in match rate, but it was the verification inside the waterfall that actually changed our bounce rate.

The checklist I'd send back to 2021 me

If you're a founder wiring up an AI agent to find emails at scale, here's the short version. It's not exhaustive, and it's biased toward B2B SaaS, which is where my experience lives.

  • Never buy verified emails as a standalone SKU. Buy them inside a workflow that verifies again.
  • Log provenance on every record. If you can't name the source, don't send.
  • Verify inside your enrichment call, not after the fact. Otherwise you're paying for the garbage.
  • Keep a human reviewer on every personalized batch. The model writes fluently; it doesn't know what's true.
  • Treat the 0.30% complaint threshold from Google's Feb 2024 bulk-sender rules as a hard ceiling. Not a target. A ceiling.
  • Watch domain reputation the way you watch MRR. It's a leading indicator for pipeline.
  • Don't let the AI write facts you can't verify. Use it for tone and structure. Verify the substance yourself.

What I still don't know

I'm not a lawyer, so I can't tell you what "GDPR-compliant outbound" means for your specific business. What I can tell you is what our team does: we keep a written record of where every record came from, we honor opt-outs on the first ask, and we pull from public sources where we can document the retrieval method. If you're selling into the EU at any real volume, get someone who actually knows this area to look at your workflow. Don't take my word for it.

Also — my numbers come from roughly 14,000 records across three B2B SaaS companies between 2021 and 2024. If you're doing recruiting, e-commerce, or anything outside SaaS, your bounce rates and reply rates will look different. Some of what I've written here transfers. Some of it won't.

What I'm confident about is this: the AI doesn't need to be smarter. It needs better boundaries. The workflow is the product. Everything else is a feature.

Erin Watanabe

Erin Watanabe
Erin Watanabe is an independent CRM and revenue workflow analyst covering prospecting integrations, lead routing, sales pipelines, API synchronization, browser extensions, campaign attribution, and sales automation. She uses ISO/IEC 27001 control objectives while checking field mapping, sync latency, webhook reliability, duplicate rate, permission scope, error recovery, attribution consistency, and audit logs. Her systems guides help revenue operations teams connect acquisition tools, preserve trustworthy records, and evaluate whether automation reduces manual work without creating hidden data debt.