i proved the tool worked first
The first version of a small tool I built did not go well. It was supposed to look at a business's website and figure out whether the business could be reached by email, an early filter before anyone spent real time on a lead. I ran it against real, live business websites before I had actually checked whether its logic was any good. It was wrong in ways I did not catch until after the fact.
I rebuilt it, and this time I did something different. Before the rebuilt version ever touched a live website, I made it prove itself against cases where I already knew the correct answer. That decision was the right one, but it did not happen as cleanly as I would like to tell it. Proving the tool worked took more than one round, not one dramatic fix.
the mistake that started this
The first mistake was simple. I let the tool run against real businesses before I had confirmed it was actually accurate. It looked for an email address on each website and reported whether the business could be reached. What I found afterward was that a meaningful chunk of what it reported as a working email address was not one at all. Some of it was automatically generated error codes that happened to look like an email format. Some of it was placeholder text that a website template leaves behind when nobody filled it in. The tool counted all of that as a real answer, which inflated how well it looked like it was working.
It also missed real information. Some websites load their content after the page opens rather than having everything there from the start, and the tool could not see any of that. On those sites, a real, working email address could sit right there on the page and the tool would report nothing.
what a real test actually requires
The fix I put in place was not complicated, but I had been skipping it. Before any tool like this touches something real for the first time, it has to be checked against a set of cases where the correct answer is already known. I write down the rules the tool is supposed to follow, hand-label a small set of real examples with the correct answer for each one, build the tool against that plan, and test it against the labeled set before it goes anywhere near live data.
When I rebuilt this particular tool, I went a step further. I had reviewers who did not build the tool hand-label the known-answer set, working independently of each other, before the rebuilt tool was ever scored against it. That mattered, because I don't think a test graded by the same hands that built the thing being tested is a real test.
it did not pass clean the first time
I would like to say the rebuilt version passed the test right away. It did not. On the first pass, it missed at least one case, a business that could genuinely be reached through a contact form on its website, which the tool did not recognize as a valid way to reach the business. The cause was a small bug in how it counted the parts of that form. I found it, fixed it, and ran the tool against the full known-answer set again. That time, it matched every case.
That was not the end of it, either. In a separate, more adversarial round of review, someone looked at cases outside the known-answer set entirely and found more problems that the first round of testing had not surfaced. Those got fixed too, in the rounds that followed, before the tool was ever pointed at the real list it was built for. Proving it worked took more than one pass. It took several.
what this changed about how I build tools
I want to move fast. When I build something, I want to see it work on the real thing right away, and that instinct is exactly what got the first version into trouble. What I learned from doing it the slower way is that the extra rounds are not wasted time. Each one caught something the tool would have gotten wrong quietly, without anyone noticing until later, which is worse than catching it up front.
I still have to fight the urge to skip the known-answer test when I am confident a tool is ready. I am getting better at not skipping it. Proving something works, on cases I already understand, before I trust it with something real, is one of the disciplines I am still building.