Experiment
The question that found two doors
Two holes in two days, in a product that sells email security and in the oldest code in the company, both found by one question that feature reviews never ask. The fixes were textbook. The finding was not.
Executive summary
On 10 October a security review of Blankitt’s DMARC service found that the address every customer publishes in DNS would accept a report for any domain and add it to the customer’s account. On 11 October a threat model of the whole company found that the sign-up path, the oldest and most reviewed code we have, would accept a tenant identifier as if it were an invitation. Both were fixed the same day they were found. Neither had been found by any review that asked “is this change safe”, because the first was built in a hurry and never re-read, and the second had not changed in a long time. Both were found by a review that started from a different question: who can reach this, and what can they make it do. The fixes are in every textbook. The question is the part worth copying.
The situation
Blankitt is a founder-led UK software company with one engineer, who is also the founder. Everything between “engineer” and “company owner” is unstaffed, which is what the operating-team project exists to fix: six role agents committed to the company’s repository, one of them a security reviewer (ADR 1). The DMARC product had been live since June with no paying customers. The decision at the start of October was to stop building features and start selling, and the first job before any stranger touched the funnel was a security review.
Two things were true that matter for what follows. The inbound report path had been written in June as the top launch blocker, in a hurry, and had not been looked at since. And a written security document existed from early 2025, correct when written, describing a company that had since added an inbound email address, a public privacy-request form, marketing forms and a dozen token links that travel in URLs.
What we thought
The security reviewer’s first job would be a formality: a tidy codebase, one engineer who knows every line, tenant isolation already designed in. The expected output was a short findings list and a long “verified OK” section. The review tooling already in use, a diff-scoped security check in the coding agent and a secret scanner on every commit, had been passing clean for months.
What actually happened
The first door: an address that is public by design
A DMARC aggregate report is sent to the address in a domain’s _dmarc record. Blankitt gives each customer an address of the form <token>@rua.blankitt.com, and the customer publishes it in public DNS. That is how the protocol works; the address cannot be secret.
The ingest code mapped the token to the tenant and then trusted the report. Whatever domain the XML said it was about, the code created that domain in the tenant’s account. There was no check that the tenant had asked to monitor it, no cap, and the attachment was written to storage before it was validated. The tier cap on domains was enforced only on the manual “add domain” route.
Anyone who could read a customer’s DNS could therefore fill that customer’s account with junk domains until the plan cap blocked real ones, forge aggregate data for the customer’s real domains, and put their own HTML into the alert emails Blankitt sends, because the domain name was placed in those emails unescaped. The generic diff review had never seen it, because the diff that introduced it was four months old and had looked fine on its own.
The fix took an afternoon. A report is now accepted only for a domain the tenant has already added; anything else is recorded as failed with the reason. The policy domain is validated against the same rule as manual entry. Decompression is capped at 30 MB. Each tenant has a daily ceiling on attachments. The demo workspace’s address rejects mail. Nine new tests, three deploys through the guarded script, health check green. Field note, 10 October.
The second door: an invitation nobody checked
The next day the reviewer was taken apart into a public pack of skills, with the company specifics moved into two files the skills read (ADR 2). The first thing the new threat-model skill was pointed at was not one product but the whole company: eight services, seven front ends, roughly a hundred and twenty ways in.
It read the 2025 security document and kept what was still true. Then it walked every entry point asking the one question. The largest finding was not in any of the new surfaces. It was in sign-up. A branch meant for invited teammates accepted the invitation token without ever checking it against a stored invitation. A tenant identifier was enough, and several public endpoints hand tenant identifiers out, because they were never designed to be secret. Anyone with one could join that customer’s account as a member, and a member could then alter billing and configure integrations, because those routes checked membership rather than role.
It had sat there through every feature review since the invite flow was written. Feature reviews ask whether a change is safe. This code had not changed.
The fix, again, was the same day. The invite token is hashed and checked against a stored invite that is unaccepted, unexpired, in an active tenant, and bound to the email that was invited; every failure returns the same generic response. Token expiry on the platform’s own session tokens, which the same pass found was not enforced, now is. Field note, 11 October.
What the two have in common
Neither fix is interesting. Validate the identifier against a record you control; never let a published value stand in for a credential; cap and validate before you write. Every one of those is in the first chapter of any application-security guide.
What they have in common is where they were found. One in code written fast and never re-read. One in code so old and so reviewed that nobody thought to read it again. A review scoped to “what changed” reaches neither.
Lessons learned
- “Is this change safe” and “who can reach this” are different questions, and only the second finds holes in code that did not change. Evidence: both doors above.
- A threat model that is not re-read is a history document. The 2025 model was correct and useless, because the surfaces that mattered arrived after it. The model now lives in the repository, is refreshed by a skill, and the review skill writes a “model drift” section whenever it finds an entry point the model does not list.
- A published token is a lookup key, not a credential. The DMARC address and the tenant identifier were both treated as proof of who was calling. Neither can be secret, so neither can be trusted alone.
- Read-only review is a feature, not a limitation. The reviewer found both doors and changed nothing; the fixes were made by a person with the finding in front of them. Nobody has to wonder what the agent altered.
Architecture changes
The code changes were small and are listed above. The change that matters is to process, and it has three edges against what most teams run:
- Against diff-scoped review (a code-scanning action, a PR bot, the security review built into a coding agent): those ask about the change. The threat model is read on every change, so unchanged code is asked the question again whenever the surface around it moves.
- Against the one-off threat-modelling workshop: the model is a file,
docs/security/THREAT-MODEL.md, with a date and a commit, maintained the way code is, with a drift section that tells you when it has gone stale. - Against the agent-fixes-it pattern: every skill in the pack reads and reports. The only files a skill writes are its own reports. The founder is the only actor with side effects, which is the governance model of the whole project.
The skills are public, MIT, in the same shape as the most-installed skills pack for coding agents, which has no security step: security-skills. The two company-specific files they read are not.
Next steps
The threat model lists further findings that are not yet closed. They are not described here, and will be, with their fixes, once they are. The review skill now runs against every change with the model as input; the test of the whole approach is whether the next hole is found by it before a stranger finds it by accident. The number that will be reported: findings opened by the review versus findings reported from outside, by quarter.
Spend in October 2026: £45.00. Spend to date: £45.00. As of 11 October 2026. Cost tracker.