Lukas Janda github hey@lukasjanda.com

// 2026-10-08

Eligibility does not belong in a prompt

A model read a hard geographic restriction as a note about timezones. Four times. The fix was not a better prompt.

A scoring system I run rated a job posting 8 out of 10 and gave this as its reason:

> Latin America CET-overlap

The posting said "based in Latin America" twice. I live in Czechia. The model had read a hard geographic exclusion — a rule about where you may *live* — as a soft note about what hours you work.

I did the normal thing first

I added a rule to the prompt. A clear one, with examples:

> A REGION LOCK that excludes the Czech Republic caps both scores at 2, however good the rest is. This is NOT the same as a timezone preference.

It did not hold. Within a week it had done it twice more:

- Remote (US and Canada) — scored 7, reasoned as "US timezone penalty" - candidates based in Colombia or Argentina only — scored 7, reasoned as "Argentina timezone workable"

Four identical failures. The model was not being careless; it was doing what models do with a rule stated in prose among twenty other rules stated in prose. "Remote" is a strong positive signal and it was reading that first.

Then it failed in the other direction

Here is the part I did not expect. Once I moved the check into code and *left the prompt rule in as well*, the model started inventing restrictions. It flagged a posting that explicitly said "Based in the Americas, Brazil, India, UK, or EU" as region-locked and capped it at 2.

So the same instruction produced false negatives and then false positives. Having two authorities on one question is worse than having the wrong one.

What actually worked

Eligibility is binary. You can take the job or you cannot. That is not a judgement call, and asking for judgement on it was the mistake.

It is now about thirty lines of pattern matching with the real failing postings as test fixtures, and every mention of region, residency and work authorisation is gone from the prompt. The prompt says only this:

> WHERE HE MAY LIVE IS NOT YOUR CALL. That is decided in code before you see the result, and judging it yourself has produced wrong answers in both directions. Score the work.

The model judges whether the work is any good. Code decides whether I am allowed to do it.

A second lesson, a day later

The deterministic check then missed REMOTE (US). My pattern required a scope of at least three characters, so it caught "US and Canada" and not "US" — the shortest and most common form of the exact thing it existed to stop.

Moving the rule into code was right. But a deterministic rule is only as good as the shapes you tested it against, and I had tested the shape I imagined rather than the shape the world uses. Both bugs were found the same way: by noticing a wrong score in real output, not by re-reading the code.