For a dev-led company, this shift matters more than it first appears, because the levers that move AI visibility are largely technical ones: structured data, feed accuracy, entity identity, third-party evidence. The discovery layer has quietly become a product surface. This article gives software teams a practical way to measure it - no marketing platform required, no budget line, just an afternoon a month and a spreadsheet.

What actually changed

The behavior shift is well documented at this point. Industry research found that 45% of consumers used AI tools for recommendations in the past year, up from 6% the year before (BrightLocal, 2026).

In B2B software the pattern is the same shape with slower public numbers: buyers ask an AI assistant for tool comparisons, "best of" lists, and stack advice, and the assistant answers with a shortlist assembled from whatever evidence it trusts. The decision has moved upstream of the click and therefore upstream of every analytics dashboard your team currently reads.

Two implications follow. First, the classic funnel instrumentation is blind to the moment where your market is increasingly decided. Rankings, traffic, conversion rates - all of it describes what happens after a click that may never come. Second, the inputs AI engines use to assemble their answers are disproportionately technical: parseable structured data, consistent product identity across the web, accurate listings, third-party mentions. Which means the people best positioned to fix them are often already on your team.

Why this is an engineering-adjacent problem

It's worth being precise about what AI engines are doing, because the mental model most teams have is wrong. An AI engine answering "what's the best error-tracking tool for a Python shop?" doesn't fetch your homepage and read it fresh. It synthesizes from evidence it has already absorbed - your structured data, your documentation's crawlability, directory listings, review-site mentions, comparison pages, third-party coverage. That evidence base is exactly the layer engineers influence most directly: schema markup, feed generation, canonical identity across properties, sitemap hygiene.

And the failure mode is engineering-flavored too. When an AI engine describes your product wrong - outdated pricing, a deprecated feature, a competitor's integration attributed to you - the cause is almost always stale or inconsistent upstream data, not malice or randomness. The engine is faithfully summarizing a messy evidence base. Garbage in, summarized garbage out. Fixing the input is a data-quality task, and data quality is a discipline your team already has.

The framework: a prompt set and a per-run log

The measurement itself is deliberately simple. It has to be, or it won't survive contact with a real roadmap.

Step one: build the prompt set. Write down twenty to fifty real questions a buyer would ask about your category. Not branded questions - nobody asks an AI engine about you by name until they already know you. Category questions: "best [category] tool for [use case]," "[your product] vs [competitor]," "what should I use for [specific job-to-be-done]," "alternatives to [competitor] for [constraint]." If your team has sales call recordings or support tickets, mine them. The questions real buyers ask are rarely the ones marketing assumed.

Step two: run and log. Run the prompt set on each major engine monthly - ChatGPT, Gemini, Perplexity, Google AI Mode at minimum. For every run, log five fields: context (date, engine, prompt, settings), presence tier (mentioned, cited, or recommended - three genuinely different outcomes), position within the answer, accuracy (does the AI's description of your product, pricing, and differentiators match your public facts), and sources (which URLs the engine cited). Five fields per run. That's the entire instrument - deliberately minimal, so it actually gets maintained.

Step three: read the divergence. The result that surprises every team on first contact: the same product is routinely recommended on one engine and invisible on another for identical questions. This isn't noise. It's the most useful signal in the dataset. Divergence between engines points at entity-signal asymmetry - your evidence base is stronger on one engine's sources than another's and entity-signal asymmetry is fixable, once you can see it. In our cross-platform AI visibility research, this pattern was the rule rather than the exception: persistent candidates, variable positions, materially different shortlists per engine for identical prompts.

A worked example

To make this concrete, consider a fictional but typical case: a small team ships an API-monitoring tool. Their twenty-prompt set includes questions like "best API monitoring tools," "API observability for small teams," and "[competitor] alternatives."

First run, three engines. Engine A recommends them in 7 of 20 answers. Engine C mentions them in 4 but recommends in only 1. Engine B doesn't surface them at all in 20. The log's accuracy column shows a recurring error on engines A and C: the AI describes a pricing tier that was retired nine months ago. The sources column shows why: both engines cite a 2025 comparison article and an outdated directory listing as their main evidence.

The diagnosis writes itself. The divergence between A and B isn't a content problem - the team's blog is fine - it's an evidence-base problem: engine B's synthesis leans on sources where the product barely exists. And the pricing error will keep propagating until the upstream sources are corrected, because the engine isn't wrong in its own terms; it's faithfully summarizing stale inputs.

The fix list, in order: correct the directory listing and the comparison article's data where possible; ship updated structured data with current pricing; add the current pricing tier to the docs page the engines can actually crawl; publish one accurate third-party comparison or integration walkthrough to seed the weaker engine's evidence base. Re-run the prompt set in thirty days. The log will tell you what moved.

That's the whole loop: measure, diagnose, correct at the source, re-measure. No platform spend, no agency, no growth hacks. An afternoon a month and a discipline your team already understands - applied to a surface it hasn't been treating as its own.

What to do with the numbers

Monthly, turn the log into three trend lines: share-of-shortlist per engine (in what percentage of category prompts are you recommended), position stability, and accuracy rate. Report them next to your existing product metrics, not instead of them - the point is that AI visibility is a product surface with product-style instrumentation, and it deserves a line in the same review. Watch trends rather than snapshots; one run is an anecdote, six runs are a signal. And keep the accuracy line honest: a wrong-fact recommendation is worse than no recommendation, because it converts curiosity into a sales call spent correcting the record.

The bottom line

Dev-led companies have a structural advantage in this shift that most haven't claimed: the levers are technical, the discipline is data quality, and the tooling is free. What most teams lack isn't capability - it's the habit of treating AI answers as a measurable product surface rather than a marketing mystery. Start with twenty prompts and a spreadsheet. The first run will produce at least one finding that surprises you. The sixth run will show you whether you fixed it. That's a measurement loop your team already knows how to run - pointed at the surface where your next customers are increasingly deciding.