One store, and one derivative of it
The index holds nothing the directory does not. It is a shape built for lookup — tokens pointing at identifiers — and it is correct only for as long as it matches the rows it was built from. When the two disagree, the index is the wrong one, every time, and the fix is a rebuild rather than an edit.
If the index is carried
If the index is rebuilt
A derivative that can be regenerated is never a thing to keep a copy of. The restore drill on the change-windows sheet rebuilds the index as its last step, and a drill that cannot rebuild it has found a defect worth having.
What goes into a token list, and what is kept out
Display name, registered name, the full category path, the post town, the outward postcode. Five fields, and no sixth without a sheet saying why.
Normalising happens once, in one place, and the same code runs on both sides of the diagram on the home page. Case folds, accents are stripped, an ampersand expands to the word, and remaining punctuation becomes a separator.
The contact route, the state, the identifier, every stamp, and every field from the listing body.
The contact route is kept out so the index cannot be mined for addresses by searching for fragments of them. The state is kept out because an unpublished row should not be findable by any term at all: unpublished entries are not indexed, rather than indexed and filtered at the end.
Running one set of rules on the stored side and another on the query side is the classic way to build a search box that cannot find a row a reader can see with their own eyes.
Query modes, and the inputs ranking is allowed to read
- All termsthe default
- Every term appears somewhere in a row’s token list.
- always
- Exact phraseasked for
- The terms adjacent, and in the order they were typed.
- always
- Any terma widening
- One term is enough, and the mode is labelled as a widening rather than presented as the result of what was asked.
- after an empty all-terms query
Ranking reads four inputs, in this order: whether the match was on a name or only on a category path; how much of the token list the match covered; whether the row’s primary category matches a category the reader is currently inside; and the alphabetical order of the display name as the final tie-break. An alphabetical tie-break is chosen deliberately — it is boring, it is stable, and a reader can predict it.
Never a ranking input: position bought from the portal, anything about an account’s payments, recency of editing, and how often a row has been returned before. A result order that rewards activity rewards whoever edits most, not whoever answers the query.
An empty result set is a surface in its own right and has to carry three things: the query as it was understood after normalising, the category tree so the reader can walk instead of type, and the nearest slug in the register when the query looks like an address that was mistyped.