Market Guy
HomeOddsDaily RecapLearnPricing
Log inSign up
Market Guy

Evidence-based prediction market research — Polymarket & Kalshi, no hype.

Product

  • Live Markets
  • Odds
  • Daily Recap
  • Daily Challenge
  • Pricing

Learn

  • All guides
  • Research library
  • Learning paths
  • Glossary
  • How markets work
  • Reading the odds
  • Smart money

How it works

  • Methodology
  • Model card
  • Data sources
  • Track record

Company

  • Account
  • Insights
  • Privacy

Market Guy · Research tooling, not investment advice.

Model card

Which model writes an analysis, how it is configured, and where it fails.

Last updated Sep 1, 2026

Model family

Analyses are written by a commercial large language model with live web search — a model that retrieves current sources at request time rather than answering from training data alone. Reports are generated; no person reviews each one before it appears.

Higher plans use a stronger, reasoning-oriented model of the same kind. The research method, the output structure and the citation requirement are identical across plans — what differs is the model doing the reading, not the standard it is held to.

Configuration

The model runs at low randomness: the task is assembling evidence, not producing variation, and the same evidence should give the same reading twice.

The search deliberately runs without a hard recency window. A filter that only admits the last few days would block exactly the primary sources a resolution analysis depends on — rulebooks, filings and schedules are often years old and still authoritative. Recency is asked for in the instructions instead, so recent developments are sought without discarding the durable sources.

The current date is given to the model on every request. A model has no clock, and without that anchor it cannot judge whether a search result is from this week or last year.

Output format

Every report has to come back in one fixed structure — assessment, sourced key facts, the case for and against, a probability, a confidence rating, the dates that matter next, and the resolution analysis. The structure is required, not requested: a response that does not match it is rejected rather than passed through. That is what stops a report from quietly arriving without its citations.

Each analysis is recorded together with the probability it stated and when it stated it. A probability that was not written down when it was made cannot be scored afterwards — that record is the reason a track record can be measured later rather than merely claimed.

Known failure modes

These are the failures we expect, listed so you can recognise them rather than be surprised by them:

  • Wrong citation index — a fact attributed to a source that does not support it. The sources are rendered visibly so this is checkable; that is the mitigation, and it is not a guarantee.
  • Confident wording on thin evidence. The confidence rating is the signal to watch, and a low rating on a precise-sounding number should be read as the warning it is.
  • Resolution rules that are genuinely ambiguous. The ambiguity rating exists for this, but a rule set can be unclear in a way no reading resolves.
  • Cross-venue matching. Where a report compares Polymarket and Kalshi, the match is made by title similarity, so two contracts about the same event can differ in their terms. A price gap between imperfectly matched contracts is not automatically an opportunity.
  • Sparse wallet history. Where a market has too few qualified wallets, the smart-money verdict withholds a signal instead of producing a weak one — an absent verdict is a result, not a failure.
  • Stale context. Underlying data is cached; a fast-moving spot price can be minutes old.

Changes

When the model or the configuration behind these reports changes, this page changes with it and the date at the top moves. That is the point of dating it.

More on how this works

  • How we research a market — The steps between a live contract and a probability estimate, and the limits of each one.
  • Data sources — Where every number on this site comes from, how often it refreshes, and what happens when a source fails.
  • Forecast track record — How the signals have actually performed — including where the sample is still too small to say.