What we decided AI would never do

Posted by

I meet a lot of engineering leaders running large teams, and there are two questions I end up asking almost all of them. How are you embracing AI-native engineering? And how is your own AI product holding up in production?

No one has a standardized answer yet. So I usually push a bit further: what does your AI system not do?

That question matters. An AI product that does everything hasn’t been designed well. It’s been deployed, and the team is hoping for the best.

The decision about where AI stops and business logic kicks in and importantly, when a human needs to step in – is usually the most important decision a company makes about its own AI product. 

Most companies haven’t made that decision. CXOs have told teams to use AI to speed up delivery, and are asking those same teams when their AI product will launch. Nobody’s asked what it’s designed to refuse to do.

We made that decision early, and we’ve held it for a long time.

AI in our interviews handles pattern recognition at scale. It asks the right follow-up questions based on the direction we give it. It keeps interview complexity consistent across every candidate, so evaluation doesn’t depend on which interviewer showed up that day. It structures signal that used to live only in someone’s private judgment. It does the repetitive, high-volume work that a human doing it manually would eventually do worse, out of fatigue if nothing else.

AI does not run the interview on its own. Every five minutes of a candidate’s interaction with our AI interviewer is shaped by business logic – rules we’ve built from being in hiring and assessment for over ten years, not from the model improvising in the moment.

Here’s why that distinction matters. A hiring decision carries consequences a model has no stake in. Someone doesn’t get an offer. Someone gets a career-defining opportunity. Someone’s team inherits a colleague they’ll work alongside for years. An AI system built on business logic, not just an LLM, can produce a well-calibrated interview and evaluation.

An LLM left to improvise cannot be held to that same standard, because nobody designed what it’s not allowed to do.

This is not a limitation we’re waiting to engineer away. It’s the actual design.

We spent years before GenAI existed manually reviewing millions of lines of code to understand what good engineering actually looks like across contexts. That work taught us something easy to lose sight of now that language models can produce fluent, confident answers to almost anything. Fluency is not the same as understanding. Confidence is not the same as correctness.

The companies that earn trust in this category over the next decade won’t be the ones with the most autonomous systems. They’ll be the ones who can explain, clearly and specifically, what their AI is not allowed to decide, and why.

Leave a Reply

Your email address will not be published.