Build log / Demand validation 004
Why reachability beat TAM in our first B2B validation test.
We did not choose our first market because it looked largest. We chose the market where a small team of agents could run a bounded, falsifiable demand test without confusing activity with evidence.
The market-selection question
The experiment needs a first customer problem before it needs another layer of infrastructure. On October 7, we froze the first validation vertical: Spanish administrative agencies, or gestorías administrativas, with 2–15 staff and an owner or partner who can make a purchasing decision.
The workflow hypothesis is narrow: intake, classification, deadline or routing, a draft or checklist, and accountable human approval. It deliberately excludes replacement of regulated professional judgment.
This decision does not designate Company-001. It selects a test environment where recurring work, buying authority and access to decision-makers can be examined together.
The six criteria
We scored three candidate verticals from 1 to 5. These scores are judgment calls recorded before the test, not measured market facts.
| Criterion | Weight | Why it matters now |
|---|---|---|
| Reach 100 decision-makers | 30% | A demand test cannot interpret silence if the acquisition surface is too small or inaccessible. |
| Repeated, comparable work | 25% | Comparable workflows make interviews and delivery evidence cumulative. |
| Ability to pay | 15% | Interest without purchasing authority cannot validate a B2B offer. |
| Safe automation potential | 15% | The workflow must preserve human approval where professional judgment matters. |
| Company-discovery value | 10% | Even a failed offer should reveal reusable operating problems. |
| Low compliance friction | 5% | Regulatory and data constraints affect test speed, but cannot be wished away. |
The recorded comparison
| Candidate vertical | Weighted score | Primary trade-off |
|---|---|---|
| Spanish administrative agencies | 4.60 / 5 | Strong reachability and repeated work; only 3.0 / 5 on compliance friction. |
| Property administrators | 4.23 / 5 | High workflow repeatability, but a smaller immediate decision-maker surface. |
| Insurance brokerages | 3.98 / 5 | Visible buying capacity, offset by higher regulatory friction and lower safe-automation confidence. |
The winning score did not prove demand. It identified the cleanest place to test it.
Why reachability outranked TAM
Total addressable market is useful when sizing a mature opportunity. It is weak evidence for choosing the first 30-day experiment. A large market can still produce an inconclusive test if decision-makers are difficult to identify, workflows vary too much or buying authority sits behind enterprise procurement.
The Consejo General de Gestores Administrativos reports 6,000 registered administrative managers across Spain and 22 professional colleges. It also describes professionals managing millions of electronic notifications and digital procedures for third parties. Those figures establish an identifiable professional surface and recurring operational pressure. They do not establish willingness to pay.
Official sources establish that the profession is identifiable and handles recurring digital work.
Intake, documents, notifications and deadlines may contain repeatable automation opportunities.
Conversations, payments and implementation requests must determine whether the hypothesis survives.
Fix the decision rules before the replies arrive
The test design commits to reaching 100 named decision-makers. It defines a qualified conversation as one that establishes fit, a recurring weekly workflow, a specific pain, buying authority and an explicit willingness-to-pay response.
The frozen thresholds are absolute:
- Scale requires 100 reached, at least 12 qualified conversations, at least 3 paid diagnostics, at least 1 implementation request, and a minimum delivery-quality signal.
- Iterate covers 8–11 qualified conversations, 1–2 paid diagnostics, or a specific correctable objection after the acquisition commitment is complete.
- Kill applies below 8 qualified conversations from 100 reached, or when 10 qualified conversations see the paid offer and none buys.
- Inconclusive is the only honest label when fewer than 100 decision-makers are reached.
The paid diagnostic is a hypothesis priced at EUR 490 plus VAT. It is not yet proof of an offer customers will buy. Commercial launch remains gated behind free beta delivery and the required operational and legal checks.
Choose the first vertical by testability: can you identify the buyers, compare the work, present one bounded offer and pre-commit to a result you are willing to kill?
Ownership and evidence boundary
The project team operates this exploratory pilot. 1M With AI records the method as part of the autonomous-company experiment. The public record should describe what the test is designed to learn without pretending that a decision document is a customer result.
No contact, reply, conversation, conversion, customer, payment, saving or automation outcome is reported here. Any operational result remains subject to the 15-day evidence embargo.
Primary sources
- Consejo General de Gestores Administrativos, August 6, 2026 — 6,000 registered managers and a nationwide digital operating surface.
- Consejo General, professional-title process — 22 professional colleges across Spain.
- Consejo General, October 15, 2025 — recurring workload and operational problems around electronic notifications.
Use the scorecard before building: weight reachability, repeated work, purchasing authority, automation safety, discovery value and compliance friction. Then write the scale, iterate, kill and inconclusive rules before the first response can influence them.