Updated

Disclosure

AI & Predictions

What a language model is allowed to do here, what it is never allowed to do — and exactly which paragraphs on this site it wrote.

A language model writes three kinds of paragraph here. It predicts nothing.

Every number on this site was computed before any of those paragraphs existed.

Three pieces of prose on BCCFC are written by a language model, and this is the complete list: the preview on a match page, the opening paragraph of a competition page, and the opening paragraph of the home page. Nothing else. Every other sentence is either fixed template copy — the same wording on every page of its kind — or a number pulled out of the database and formatted. Where a sentence looks tailored to one fixture, what is tailored is the numbers inside it.

Those three paragraphs are written from a closed fact sheet: the same probabilities, scoreline, form string, head-to-head record and confidence grade the page renders around them, and nothing else. The model is not connected to the database, cannot look anything up, and is never shown the internet. If a fact is not on the sheet, there is no route by which it can reach the page — which is why a preview here will not tell you who is injured or what a manager said.

One more rule is specific to this site rather than to the platform: no bookmaker is named anywhere on BCCFC, and that applies to generated copy too. Text that comes back with a bookmaker's name in it is discarded rather than published, and the page keeps whatever it had before.

We are saying so on its own page because “AI football predictions” is a phrase that now means almost nothing. It is used for a fitted statistical model, for a language model writing tips, and for a spreadsheet with a logo on it. Those are three different products with three different failure modes, and a reader deserves to know which one they are reading. Here the model writes; it does not predict.

The rules that apply, now that it is on

Written down before it was switched on. Unchanged since — only the tense is.

  • A language model may only describe numbers that already exist. It is handed the same prediction data the page renders, and it may explain it. It may not produce a probability, a scoreline, an injury, a quote or a result.
  • If the words and the numbers ever disagree, the numbers are right and the sentence is a defect. Never the other way round — no table on this site is edited to agree with a paragraph.
  • Generated copy is generated once and stored, tied to the facts it was written from. When those facts change, the text is regenerated rather than left to describe a fixture that no longer exists.
  • Nothing about it changes the model. Predictions are computed before any text is written and are not aware that any text exists.

So what is doing the predicting?

A statistical model, fitted to results and expected goals: rolling team ratings feeding a Dixon-Coles adjusted Poisson scoreline grid, blended with the de-vigged market price on 1X2. It is not a neural network and we are not going to call it one. It has a fixed number of parameters, all of which are written out on the mathematics page, and a published track record you can check against a baseline.

The reason to prefer that to a language model here is narrow and practical: a language model has no way to be calibrated. You cannot ask it whether the things it called 60% happened 60% of the time, because it has no such quantity. Our model does, and that number is on the accuracy page whether it flatters us or not.