Skip to main content

    Deep dive

    Why this isn't a deep-research prompt

    Anyone can generate a detailed report about a market. Almost nobody can generate the same report twice. This page is for people who want to know what sits between those two sentences - the part that does not show up in the output, but decides whether the output is worth building a positioning decision on.

    What a strong deep-research run actually gives you

    Honestly: quite a lot. A Pro-tier deep research run, or an agentic workflow wired together over a long weekend, will surface real competitors, quote real copy, and produce something that reads like a category overview. For orientation, that is genuinely useful.

    What it does not give you is a foundation. The output is shaped by the prompt rather than by the category. It is not reproducible - the second run disagrees with the first in ways you cannot attribute to the market. Coverage of non-web surfaces is thin, so brand behaviour that lives in social and paid channels is largely invisible. And the synthesis layer is where confidence exceeds evidence, which is precisely the layer a strategist has to defend.

    The gap is not intelligence. The models are extremely capable. The gap is everything that has to be decided before and around the model call.

    The five hard problems

    Each of these is a problem before it is a feature. Described here at the level of architecture - the rubrics, weights, and context layer stay ours.

    01

    Scope resolution

    A category is not a keyword. Ask a model for "competitors in electric bikes" and you get whatever the training data and the first ten search results agree on.

    Scope is resolved before any research runs: the category is defined from the client brand's own perspective - what it actually competes for, in which geography, at which layer of the market. Everything downstream inherits that frame, which is why two runs on the same brand stay comparable.

    02

    Source acquisition and classification

    Most tools read whatever a search index returns and treat every page as equally informative. A press release then reads as a capability, and a careers page reads as a strategy.

    Sources are pulled from the competitor's own surfaces, the wider web, and social and paid channels - then typed and weighted by what each surface can legitimately tell you. Claims about positioning and claims about delivery are drawn from different classes of evidence, on purpose.

    03

    Evidence-to-claim binding

    Language models synthesise. Synthesis without a citation path is the exact failure mode that makes a strategist re-verify everything before putting it in front of a client.

    Claims exist only where evidence exists. Every scored attribute and every territory statement carries a path back to the source it came from. Where the evidence is thin, the output says so instead of smoothing over the gap.

    04

    Reproducibility

    Run the same prompt twice and you get two different reports. Both plausible. That is fine for exploration and unusable as a foundation for a positioning decision.

    Reproducibility is a property of a fixed methodology and a fixed schema, not of a model's mood. The structure of the read - the dimensions, the scoring frame, the comparison logic - is held constant, so the thing that changes between runs is the market, not the narrator.

    05

    Context engineering for downstream AI

    Handing an assistant a 60-page report and asking a question is a retrieval problem dressed up as reasoning. Accuracy collapses with the size of the dump.

    Outputs are structured so a downstream assistant receives only the slice relevant to the question - the right competitors, the right dimension, the right evidence. That is what makes AI-assisted reasoning on top of this material accurate rather than confidently approximate.

    Why three days of agentic plumbing doesn't get there

    Orchestration is the easy part. Crawling, fan-out, retries, a queue, a scheduler, a citation field - a good engineer builds that in a week and it works.

    What cannot be assembled in a week is the judgement encoded underneath: eight years of framework work across 2,000+ companies, a scored schema that makes competitors comparable on dimensions that matter strategically rather than dimensions that are easy to extract, and a calibration corpus that tells us when a score is wrong. Without that layer, an agentic pipeline produces well-plumbed opinions at scale.

    The methodology is the product. The pipeline is how it runs.

    What we deliberately don't show

    Three things stay private, and we would rather say that plainly than pretend the page is exhaustive: the scoring rubrics behind each strategic dimension, the context-engineering layer that decides what a model sees at each step, and the source-weighting model that determines how much a given surface is allowed to influence a claim.

    Everything in a delivered report is traceable and checkable by the strategist using it. How the report is produced is not.

    What compounds

    Because every run uses the same schema, every run adds structured, comparable data about a category - not a document that gets filed and forgotten. Positions become time series. Territory ownership becomes something you can watch move.

    Monitoring makes those movements visible between engagements. MCP makes the whole structured read queryable from inside the AI tools a strategist already works in, with the right slice served per question rather than the whole file.

    Questions that go deeper than this page: guntis@theogrowth.com.