Skip to content
Signalforge

GTM engineering · Chicago · remote

GTM systems that compound.

We build the pipeline that replaces manual SDR and RevOps volume: enrichment, scoring, routing, outbound, signals and reporting, wired together and measurable. Every system below is a working demo, not a slide.

SDR
sends the messages. Volume from people.
RevOps
keeps the CRM honest. Process and reporting.
GTM engineering
builds the machine that does both. Volume from systems.

One sentence: GTM engineering is the discipline of building the automated systems that find, qualify, contact and route buyers, so a small team runs at a large team's volume without the headcount.

What we build

Five systems, in this order.

The order is the point. Each system consumes what the one before it produces, so building reporting before enrichment measures garbage precisely. Every card links to a working demo.
  1. 1demo →

    Enrichment pipeline

    Domain in, qualified contactable record out. Qualify before you buy the email.

    domains → scored records

    Why first: Everything downstream scores and routes what enrichment produces. Bad input, bad machine.

  2. 2demo →

    Scoring model

    Weighted timing signals with decay. Fit says who could buy; timing says who is buying now.

    records + signals → ranked list

    Why #2: Decides who gets a message this week. Timing over fit.

  3. 3demo →

    Reply routing

    Classify every inbound reply, update the CRM, page the owner. Minutes, not mornings.

    replies → CRM + Slack

    Why #3: The first system that touches revenue. Minutes matter.

  4. 4demo →

    Signal detection

    Daily watch on job boards, funding, exec moves and ad libraries. Accounts fire; lists stop going stale.

    feeds → fired accounts

    Why #4: Feeds scoring with fresh timing instead of stale lists.

  5. 5demo →

    Reporting layer

    Reply, interested, meeting and pipeline by segment × persona. The aggregate hides the two cells that carry everything.

    activity → segment × persona

    Why #5: Tells you which segment × persona is actually working, so you can double down.

Interactive demos

The systems, running.

Everything below runs in your browser on labelled sample data. No API calls, no tracking. The logic is the same shape we ship into Clay, HubSpot, Salesforce and n8n; the data is invented so nothing here is a client claim.

01Demo data

Enrichment pipeline

Enter a domain. Watch it move from a broad source through a scrape, an evidence-only description and a scored ICP test, to the threshold. Email credits are spent only after it qualifies.

client-side · no calls made
  1. Broad sourceFirmographic lookup

    Waiting.

  2. Site scrape4 pages, summarised

    Waiting.

  3. AI descriptionFrom evidence only

    Waiting.

  4. ICP scoreTests + disqualifiers

    Waiting.

  5. Threshold cutQualify at ≥ 70

    Waiting.

  6. Email waterfallOnly after qualify

    Waiting.

Output record

No record yet. Enter a domain and run the pipeline.

Try a qualified one and a disqualified one to see the waterfall gate.

Credit burn

Same batch, two orders of operations. Email waterfalls cost roughly 2.3 credits a record; a scrape and a small-model classification cost about 0.08.

1,000
22%

75% fewer credits, or 1,794 credits saved on this batch. The saving is the disqualify rate; it grows with a stricter ICP.

Illustrative credit prices. Swap for your provider's rate card.

02Demo data

Scoring model: timing, not just fit

Eight accounts, five signal types, explicit weights. Flip between ranking by fit and by timing, toggle decay on year-old signals, and see why three fresh signals beat a perfect firmographic match.

  1. #AccountFitTimingTotal

Why this rank

Kestrel

kestrelhq.com

35% × fit 72 + 65% × timing 75 = 74

  • Funding round6mo ago

    Series B, $34M

    Fresh capital funds headcount and tooling. Budget exists for roughly two quarters.

    +25
  • Hiring for buyer function18d ago

    4 SDR + 2 AE roles posted

    Open roles for the people who would use the product mean the function is growing right now.

    +30
  • New executive1mo ago

    New VP Sales (ex-Gong)

    A new leader re-evaluates the stack in their first 90 days.

    +20

The lesson

Harborstack is the best fit on paper (95/100) but has no fresh signal. Kestrel fits less well (72/100) but 3 things changed there recently, so it outranks on total score.

Fit tells you who could buy. Timing tells you who is buying this quarter. A rep working the fit-only list sends Harborstack a message nobody asked for. Decay keeps two-year-old funding rounds from pretending to be news.

03Demo data

Reply routing

Simulated replies are classified, written to the CRM and posted to the owner's channel as a webhook chain. Then drag the response-time slider to see what slow follow-up costs.

Inbound replies

Event log

0 events

No events yet.

Process a reply and the webhook chain appears here.

Nothing in #inbound-hot yet.

Time to first response

A positive reply is a hand raised. The routing above exists so a human answers in minutes, not the next morning. Drag to see what waiting costs.

30 min

Illustrative curve. Relative close probability vs. time to first human response on a positive reply.

Relative close probability

86%

of a 5-minute response

Close rate on positives

20.6%

base 24% at 5 min

Deals lost per month at 40 positive replies

1.3

Each one waits 30 min at this setting.

04Demo data

Signal detection

A daily watchboard over ATS job boards, funding, exec changes and ad libraries. Each rule has a content test and a time window; accounts that pass both fire into the scoring model.

runs daily06:00 America/Chicago · last run 2026-09-01 · 142s · 16 items checked

Rule: 2+ buyer-function roles opened in 30 days. Sources: Greenhouse-style, Lever-style, Ashby-style.

  • Harborstack2d agocontent ✓≤ 30d ✓fired

    SDR (x3) · AE (x3) · Sales Enablement Manager

    Greenhouse-style board

  • Ferro Labs5d agocontent ✓≤ 30d ✓fired

    Head of Sales Development

    Ashby-style board

  • Tessellate10d agocontent ✕≤ 30d ✓

    Senior Backend Engineer (x3)

    Greenhouse-style board

  • LumenOps12d agocontent ✓≤ 30d ✓fired

    Account Executive (x2), first sales hires

    Lever-style board

  • Kestrel18d agocontent ✓≤ 30d ✓fired

    Sales Development Representative (x4) · Enterprise AE (x2)

    Greenhouse-style board

Accounts that fired

6 accounts

Each fired account is pushed to the scoring model with the rule id and date, so the timing score above updates before anyone writes a message. Accounts that fire two rules in the same week go to a human first.

05Demo data

Reporting layer

The same 90 days of outbound, two ways. The aggregate looks fine. Split it by segment × persona and two cells are carrying sixteen.

Last 90 days · demo data

Reply rate, all segments and personas

3.9%

Sent
15,260
Replies
590
Meetings
90
Pipeline
$2.84M

This is the number in most weekly reports. It looks like a healthy program. It is actually two segment × persona pairs carrying sixteen that are losing money, and the aggregate cannot show you which.

Prompt lab

How we write prompts that hold at volume.

A prompt that works on ten companies is a demo. A prompt that gives the same answer on ten thousand, cites its evidence and refuses to invent is a system. The difference is mostly discipline.

Same task: score a company against an ICP.

ICP scoring prompt
Holds at volumeSame output shape every time. Auditable. Cheap to re-run.
You score companies against an ICP. You never invent facts.

Evidence you may use: the SOURCE block only. If a test cannot be verified from SOURCE, mark it "unknown" and award 0 points.

TESTS (award full points or 0):
1. b2b_software (20): sells software to businesses. Fail for services, agencies, hardware, consumer.
2. stage (15): Series B–D, OR 80–600 employees.
3. sales_team (20): ≥3 open or filled SDR/AE roles.
4. buyer (15): pricing or customers imply ACV ≥ $15k (enterprise tier, 'talk to sales', enterprise logos).
5. crm (10): HubSpot or Salesforce detected.
6. geo (10): HQ in US, CA, UK or EU.
7. outbound_motion (10): sequencing tool detected OR SDR roles open.

DISQUALIFIERS (any present ⇒ score = min(score, 30)):
- consumer product · agency/consultancy/staffing · <40 employees · no sales function · regulated life sciences · government

OUTPUT strict JSON:
{
  "tests": [{"id": "...", "pass": true|false|"unknown", "evidence": "<quote or paraphrase from SOURCE, ≤ 25 words>"}],
  "disqualifiers": ["..."],
  "score": <0–100>,
  "reason": "<≤ 40 words, must cite at least two specific facts>"
}

SOURCE:
{{source}}
  • Each test is binary with a written pass condition.
  • Disqualifiers cap the score, so no amount of fit rescues a bad company.
  • 'Unknown scores zero' removes invention and makes silence visible.
  • Evidence is quoted, so a human can spot-check 50 records in 10 minutes.
  • Weights are explicit; the score distribution is calculable before you run it.

Currently showing the Holds at volume version on small screens.

Symptom → fix

What a scoring prompt does wrong in production, and the change that fixes it.

Keeps bad companies

CauseNo disqualifiers, or disqualifiers phrased as negative points instead of caps.

FixAdd explicit disqualifiers that cap the score. Test against 20 known-bad records before scaling.

Drops good companies

CauseTests require evidence the scraper rarely finds (exact employee count, funding stage).

FixAllow alternative evidence per test ('Series B–D OR 80–600 employees'). Log 'unknown' rates per test and fix the scraper, not the prompt.

Generic reasons

CauseReason field is unconstrained and the model pads it.

FixCap the reason at 40 words and require at least two cited facts. Reject outputs that fail the check.

Score clustering

CauseUnanchored scale. The model hedges toward the middle.

FixReplace the scale with weighted binary tests. The score becomes arithmetic, not opinion.

Different answer on re-run

CauseVague criteria and temperature above zero.

FixBinary tests, temperature 0, and a golden set of 100 records that must score within ±5 on every deploy.

Bill grows faster than the list

CauseFrontier model reads every record, including the 80% that a keyword rule could reject.

FixCheap-model-first routing: a small model applies disqualifiers, the frontier model scores survivors only.

Cheap-model-first routing

A small model culls, the frontier model scores survivors. Same output quality on the records that matter.

  1. Raw domains100% of records

    10,000 records

  2. Rules62% of records

    Regex + firmographic filters. Free.

  3. Small model28% of records

    Applies disqualifiers. ~$0.0003 / record.

  4. Frontier model28% of records

    Full ICP scoring. ~$0.012 / record.

  5. Qualified19% of records

    Enriched and routed.

Frontier on everything

$120

Routed · 69% less

$37

Per 10,000 records, illustrative list prices. The qualified set is identical because the small model only applies disqualifiers it can verify.

Outcome patterns

What these systems tend to produce.

Industry patterns · not client claims

Automate ~80% of SDR workflow

Enrichment, scoring, first-touch drafting and reply triage run as systems. Humans handle positive replies and calls.

80–100

meetings per rep per month, reported in public patterns

  • Qualify before enrich
  • Signal-timed sequencing
  • Reply routing under 5 minutes

Caveat: Depends on a real ICP and clean domains. Volume without deliverability is just spam.

Dynamic content from signals

First lines and offers generated from the signal that fired (role opened, round closed, exec joined), not from a template.

2–3×

reply-rate lift over static templates, in public patterns

  • Signal → hook mapping
  • Persona-specific proof
  • Forbid invention in generated copy

Caveat: Lift collapses when the signal is stale. Decay rules matter.

Call language → outbound loop

Phrases buyers actually use on recorded calls are mined weekly and fed back into sequences and ICP tests.

Weekly

feedback cycle from calls to copy, instead of quarterly rewrites

  • Transcript mining
  • Objection clustering
  • Copy tests by segment

Caveat: Requires consent-compliant recording and a reporting layer that splits by segment.

Patterns drawn from public GTM engineering practice. These are illustrative, not client claims. Real results depend on list quality, offer and deliverability.

Commercial foundation

The judgement under the systems.

Automation multiplies whatever it is pointed at. These three things decide whether it multiplies pipeline or spam: how deep the ICP goes, whether the funnel maths supports the plan, and whether the mail lands.

ICP depth

Most ICPs stop at layer one. Reply rates live in layers three and four.

  1. 1

    Firmographic

    Industry, size band, stage, geography

    Table stakes. Every vendor filters on this. It removes nothing your competitors keep.

  2. 2

    Technographic

    CRM, sequencing tool, data warehouse, ATS

    Says whether integration is easy and whether the buyer has already bought the category.

  3. 3

    Situational

    Hiring, funding, exec change, expansion, contract renewals

    Timing. This is where reply rates come from.

  4. 4

    Disqualifiers

    Agencies, regulated verticals, sub-40 headcount, no sales team

    The cheapest lever. A written disqualifier list removes 40 to 60 percent of most bought lists.

Deliverability checklist

0/8

Tick what you have. Anything marked critical that is missing means the rest of this page is not worth building yet.

4 critical items missing. Fix before any campaign runs.

Funnel maths

Before building anything, check that the plan closes. Move any number and the rest recompute. Defaults are mid-range for cold outbound to a well-defined ICP. Replace with your numbers.

2
1,500
4.5%
28%
70%
45%
22%
$38K
  1. Sends3,000
  2. Replies135
  3. Positive37.8
  4. Meetings26.5
  5. Opps11.9
  6. Won2.6

Bars are square-root scaled so the small stages stay visible. Values are per month.

Pipeline created / month

$452K

Closed-won / month

$100K

26 meetings a month at 3,000 sends. Each 1-point gain in reply rate is worth $101K of monthly pipeline at these ratios.

Engagement options

Three ways to work with us.

Start with the teardown if you are unsure. It is the fastest way to find out whether your problem is list quality, timing, response speed or reporting, and each of those has a different fix.

Systems teardown

45 minutes

Live call

Heads of Growth, Sales or RevOps who suspect the machine is leaking.

  • Walk through your current stack from list to meeting
  • Find where credits, replies and hours are lost
  • Leave with a build order and the one fix to make this week

Project build

4 to 8 weeks

Fixed scope

Teams that need one of the five systems shipped and handed over.

  • One system, built in your tools (Clay, HubSpot, Salesforce, n8n, Slack)
  • Written runbook, prompts and golden test set
  • Two weeks of tuning after launch

Embedded GTM engineer

Monthly

Part-time, in your Slack

Series B+ teams that want the whole loop built and iterated over a few quarters.

  • Build order executed system by system
  • Weekly reporting by segment and persona
  • Prompt, signal and copy iteration from live results

Not sure which? Send a sample ICP.

We reply with a scored sample list, the disqualifiers you are missing, and a build order. No deck.

Send a sample ICP