On this page
Most outbound teams run disconnected tools. Sequencer over here. List vendor over there. Call recordings somewhere else.
They're sitting on a goldmine of data and treating it like three separate piles of dirt.
Here's the stack we run internally - and what we've started building for clients who want their systems to actually compound over time.
Part 1: The List Processor
Every list you buy is dirty. Dead companies. Defunct emails. Duplicates at both the contact and company level.
You're paying to send emails to ghosts.
The Merge and Dedupe Flow
When you get multiple contact lists (and you will - from conferences, intent vendors, enrichment tools), you need to consolidate them before you do anything else.
Step 1: Merge everything into one master file
Throw all your CSVs into one dataset. Don't worry about duplicates yet - just get everything in one place.
Step 2: Contact-level deduplication
Dedupe based on:
- Personal LinkedIn URL
- Email address
That's it. Those are your unique identifiers for humans.
Step 3: Company-level deduplication
This is where most teams screw up. They dedupe on company name only and end up with three records for "Acme Inc" because one says "Acme, Inc." and another says "ACME Incorporated."
Dedupe in this order:
- Company LinkedIn URL (most reliable)
- Website domain
- Google Place ID (for local businesses)
- Company name + full address (last resort)
The order matters. LinkedIn URL is clean and canonical. Company name is messy and error-prone.
Validation Layer
Once you have a clean list, you need to validate it. Two things matter:
- 1. Is the website still alive?
- Sounds obvious. Most teams skip this. They're emailing domains that 404.
- Run a simple HTTP check. Flag anything that doesn't resolve or returns 5xx errors.
- 2. What's their MX provider?
- This is the hidden killer.
Some domains run email security gateways that destroy cold email deliverability. If you're sending to companies using certain enterprise security tools, your emails are going straight to spam - or getting blocked entirely.
Security gateways that kill cold email:
| Gateway | Notes |
|---|---|
| Proofpoint | Enterprise-grade, aggressive filtering |
| Mimecast | Common in mid-market, strict reputation scoring |
| Barracuda | Blocks unknown senders aggressively |
| Cisco/IronPort | Enterprise standard, reputation-based |
| Symantec/MessageLabs | Legacy but still prevalent |
| Forcepoint | Government and finance verticals |
| FireEye | Security-conscious orgs |
| TrendMicro | APAC heavy, strict policies |
| Sophos | SMB market, cloud filtering |
| AppRiver | MSP-managed clients |
| SpamHero | Aggressive spam scoring |
Pull the MX records. Flag domains with these providers. Adjust your volume or skip them entirely.
Regex-Based Business Classification
Before you burn credits on AI enrichment, you can classify 80% of companies using pure regex against their homepage scrape.
SaaS indicators (regex matches):
- free\s+trial
- \bapi\b
- \bintegrations?\b
- per\s+month
- \bdashboard\b
- \bworkflow\b
- \bplatform\b
- \bsubscription\b
- start\s+free
api[\s._-]?docs?
\bdeveloper[s]?\b
\bsdk\b
Agency/Services indicators:
- our\s+services
- \bportfolio\b
- get\s+a?\s*quote
- \bconsulting\b
- \bagency\b
- our\s+work
- \/services\b
- \/our-services
E-commerce indicators:
- add\s+to\s+cart
- \bcheckout\b
- free\s+shipping
- buy\s+now
- \bin\s+stock\b
- \bshopping\s+cart\b
Local business indicators:
visit\s+us
store\s+hours?
\bdirections?\b
\(\d{3}\)\s*\d{3}[-\s]?\d{4} # Phone number
\bmon(?:day)?\s*(?:-|through|to)\s*fri(?:day)?\b
walk[\s-]?ins?\s+welcome
Run these patterns against nav, footer, and link text. Main content is too noisy.
Extracting Research Links
The real gold is in the links. When you scrape a homepage, extract and categorize URLs so your research agents (Claygent, Perplexity, whatever) can go directly to the right pages.
Priority link patterns:
| Category | Regex Pattern | Why It Matters |
|---|---|---|
| Careers | /careers?|jobs?|hiring|join-us|open-positions | Hiring = growth signal |
| Case Studies | /case-stud|testimonial|success-stor|\/customers?|\/reviews? | Social proof = serious company |
| Pricing | /pricing|plans|packages | Published pricing = product-led |
| Documentation | /docs?|documentation|api|developer|\/integrations? | Docs = SaaS/tech product |
| Competitors | /compare|\/vs|alternatives? | Comparison pages = market awareness |
Send your agent directly to /careers instead of asking it to "find hiring information." You'll get better results and burn fewer tokens.
Optional: Homepage Intelligence
If you want to get serious, scrape the homepage and run it through an LLM.
You'll extract:
- What they actually sell (not what their LinkedIn says)
- Recent news or announcements
- Hiring signals
- Technology indicators
This feeds into your copy. More on that later.
Part 2: The Analytics Engine
Your sequencer is generating data every day. Open rates. Reply rates. Bounce rates. Most teams look at this once a month in a dashboard.
That's not analysis. That's archaeology.
The Ingest Layer
First, you need to pull bulk analytics out of your sequencer automatically.
Whether you're on Smartlead, Instantly, Apollo, or something else - build a pipeline that extracts performance data and normalizes it into a consistent format.
Key metrics per email step:
- Sends
- Opens
- Replies
- Positive replies
- Bounces
- Unsubscribes
Key metrics per campaign:
- Total pipeline generated
- Meeting book rate
- Reply-to-meeting conversion
- Skill 1: Flag Winners and Create Split Tests
Here's where it gets interesting.
Once you have performance data flowing in, you build logic to:
Identify statistical winners
Which subject lines beat the control?
Which email steps have above-average positive reply rates?
Which campaigns outperform the benchmark?
Auto-generate split test candidates
Take your winners
Create variations that test one variable
Queue them for your next send
Most teams guess at what to test. This system tells you exactly where the opportunity is.
Skill 2: Anti-Fingerprinting Protection
The dark side of cold email at scale: ESP fingerprinting.
If you send the same copy structure across thousands of emails, Gmail and Microsoft start recognizing it. Not the words - the pattern. The rhythm. The format.
You need to systematically rotate:
- Sentence structures
- Spacing patterns
- Opening hooks
- CTA variations
The analytics engine can flag when campaigns start showing deliverability decay and suggest pattern rotations before you get flagged.
Part 3: The Conversation Processor
Your sales calls are the most undervalued asset in your entire GTM operation.
Every call contains:
- Pain points in the prospect's own words
- Objections you'll hear again
- Competitors they're evaluating
- Technology they're already using
And 99% of teams do nothing with this data.
The Pipeline
Record your calls. Push transcripts into your system. Build a processor that extracts and caches:
Pain points
What problems do they mention?
What language do they use to describe frustration?
What metrics do they care about?
Objections
What pushback comes up?
When in the conversation does it happen?
How do successful reps handle it?
Competitors
Who else are they talking to?
What do they say about the competition?
What features matter in the comparison?
Technology
What tools are they already running?
What integrations matter?
What are they migrating away from?
The Feedback Loop
Here's where the three systems connect.
Take the most common pain point from your calls. Look at your analytics engine. Are you testing copy that uses that exact language?
If not, you have a split test to run.
Take a frequent objection. Look at your sequences. Are you preemptively addressing it in step 2 or 3?
If not, you have a variation to build.
Your ICP knowledge base gets smarter every week. Your copy gets more resonant. Your targeting gets tighter.
The Compound Effect
Most GTM stacks are static. You set them up, run campaigns, and hope for the best.
This stack learns.
Clean lists mean better deliverability. Better analytics mean smarter tests. Call intelligence means copy that sounds like your prospects wrote it.
Each piece feeds the others. The system gets sharper over time.
Implementation Notes
You don't need to build this all at once.
Start with the list processor. Every percentage point of bad data you remove improves everything downstream.
Add analytics ingestion next. Even basic winner flagging will identify opportunities you're currently missing.
Add call processing last. This is highest value but also highest complexity.
The order matters because each layer improves the signal-to-noise ratio for the next.
Now go build it.
Originally published on X: 3 Claude Code GTM skills you can stand up in an hour that will make you superhuman
