Every accommodation booking system eventually runs into the same wall. Owners describe their prices in prose, and software needs rows in a table.
What arrives looks like this, give or take a language. The colours show which sentence becomes which kind of row:
| Row | Captures |
|---|---|
| season | Date range, nightly rate, guests included, extra guest fee, cleaning fee and when it is waived |
| holiday | Fixed night count and a package total |
| pet fee | Per-animal rate, maximum animals, maximum weight |
| discount | Type night_bonus: 7 nights required, 1 night free |
| rule | Arrival and departure day, minimum stay |
Six sentences become five rows, with every ambiguity resolved on the way. The model returns this:
{
"property": { "pricing_year": 2026 },
"seasons": [
{
"season_name": "High season",
"days_of_week": "1,2,3,4,5,6,7",
"start_date": "2026-06-15",
"end_date": "2026-08-22",
"base_price_per_night": 35000,
"base_guests_included": 6,
"extra_guest_per_night": 5000,
"cleaning_fee": 15000,
"cleaning_fee_waived_at_nights": 7,
"priority": 0,
"notes": null
}
],
"holidays": [
{
"holiday_name": "Christmas package",
"start_date": "2026-12-24",
"end_date": "2026-12-28",
"fixed_nights": 4,
"package_total_price": 180000
}
],
"pet_fees": [
{
"fee_per_animal_per_night": 2000,
"is_one_time_fee": 0,
"max_animals": 2,
"max_weight_kg": 15
}
],
"discounts": [
{
"discount_type": "night_bonus",
"days_of_week": "1,2,3,4,5,6,7",
"min_nights_required": 7,
"bonus_nights": 1
}
],
"rules": [
{
"rule_type": "specific_days",
"start_date": "2026-06-15",
"end_date": "2026-08-22",
"min_nights": 7,
"arrival_day": 6,
"departure_day": 6
}
]
}Note the two date conventions: the season’s end_date is its last night, while the holiday’s end_date is the departure day, so four nights from 24 December end on the 28th. Weekdays are ISO numbers, so Saturday is 6. Fields that do not apply are left out here.
Someone has to do that translation. For years it was a person, at maybe twenty minutes per property per year, which is fine for five properties and miserable for a hundred. This post is about handing that job to a language model, and about the parts nobody mentions: what the JSON does when the model has a bad day, how to pick a model, and what a run actually costs.
The five stages between an owner’s paragraph and a saved price row. Only stage three involves a model.
The job is extraction, not arithmetic#
The single most important design decision came first, and it is the one I would argue for hardest.
It reads text and emits structured data. A deterministic pricing engine, written in ordinary code and covered by ordinary tests, takes those rows and computes what a stay costs. If the engine is wrong, it is wrong the same way every time and a test can pin it down. If a model does the arithmetic, it is wrong occasionally, differently each run, and confidently.
There is a second, smaller use of the model in the same system: an operator preparing a quote can ask for a plain-language breakdown of a specific booking, and the model returns two or three notes about what to watch for plus a one-line calculation. That is a draft for a human who is about to write an email. The number the customer is charged still comes from the engine. The distinction sounds pedantic until the first time a model quietly adds a deposit into a total.
The output shape#
The whole approach rests on the model producing exactly one JSON object and nothing else. The prompt opens with that requirement and does not soften it:
Output ONLY a single valid JSON object. No markdown fences, no prose, no explanation, no leading or trailing text.
The schema is spelled out field by field, with types and nullability, because a model that has to guess whether a missing cleaning fee is 0 or null will guess differently on Tuesday than it did on Monday:
{
"property": {
"name": "<string from input>",
"pricing_year": 2026,
"max_capacity": 10,
"units_available": 1,
"free_child_age_limit": 3
},
"seasons": [
{
"season_name": "<unique name>",
"days_of_week": "1,2,3,4,5,6,7",
"start_date": "YYYY-MM-DD",
"end_date": "YYYY-MM-DD",
"base_price_per_night": 35000,
"base_guests_included": 6,
"extra_guest_per_night": 5000,
"cleaning_fee": 15000,
"cleaning_fee_waived_at_nights": 7,
"priority": 0,
"notes": null
}
],
"holidays": [],
"pet_fees": [],
"discounts": [],
"rules": []
}The other four arrays are specified the same way, one object shape each.
Around the schema sits about four hundred lines of rules, and they are the actual product. A few of the ones that were learned the hard way:
State every convention explicitly, especially inconsistent ones. A season’s end_date is the last night it covers. A holiday package’s end_date is the departure day, so a four-night package starting 30 December ends on 3 January. That inconsistency is inherited from the database schema and cannot be wished away, so the prompt says it twice, with a worked example, and the validator checks it afterwards.
Give the number formats no room. Prices are integer units of currency with no separators. The prompt lists the shapes that appear in real text (35.000 Ft, 35 000 Ft, 35e Ft, 35 ezer Ft) and states that they all mean 35000. Without that, thousands separators come back as decimal points roughly one time in ten, and 35,000 quietly becomes 35.
Handle the range case with a rule, not with judgement. “30,000 to 35,000 per night” is a real thing owners write. The instruction is to take the lower bound and record the range in a note. Any rule will do as long as it is fixed; what you cannot have is the model picking a different end of the range depending on how the sentence was phrased.
Say what not to extract. Deposits, cancellation policies, check-in times, payment terms, house rules and marketing copy are all in the source text and none of them are pricing. Listing them explicitly removed a whole class of invented rows.
Name the things downstream code depends on. The engine matches high season with a SQL LIKE on a specific word. So the prompt says: if there is a main season, its name must contain that exact substring, and here are acceptable and unacceptable examples. This is ugly and it is honest. The alternative is an extraction that looks perfect and silently fails to link holiday prices.
Skip what the schema cannot represent. Early bird discounts, last minute deals, group size discounts and loyalty rates all appear in real price lists, and none of them fit the discount table. The prompt names them and says to omit them. Before that, the model helpfully invented fields to hold them, and the importer dropped those rows without telling anyone.
Then a self-check section at the end lists the invariants: every cross-reference resolves, season names are unique, dates are ten characters and inside the right year, enums are from the allowed set. Asking for a check before emitting is not a guarantee, but it measurably reduces the number of objects that fail validation on arrival.
Calling it#
The call itself is about thirty lines and has no framework in it:
function callOpenRouterAPI(string $prompt, string $model,
int $maxTokens = 1024, int $timeout = 120): string {
$body = json_encode([
'model' => $model,
'max_tokens' => $maxTokens,
'stream' => false,
'messages' => [['role' => 'user', 'content' => $prompt]],
]);
$ch = curl_init();
curl_setopt_array($ch, [
CURLOPT_URL => 'https://openrouter.ai/api/v1/chat/completions',
CURLOPT_RETURNTRANSFER => true,
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => $body,
CURLOPT_HTTPHEADER => [
'Content-Type: application/json',
'Authorization: Bearer ' . OPENROUTER_API_KEY,
],
CURLOPT_TIMEOUT => $timeout,
]);
$response = curl_exec($ch);
$httpCode = curl_getinfo($ch, CURLINFO_HTTP_CODE);
curl_close($ch);
// error handling omitted
$data = json_decode($response, true);
return $data['choices'][0]['message']['content'] ?? '';
}OpenRouter is a single OpenAI-compatible endpoint in front of most commercial and open-weight models. You send a model identifier as a string and it routes the request. That is the entire value proposition, and for this kind of work it is worth a lot, because it turns “which model should we use” from an architecture question into a config value:
OPENROUTER_MODEL_PRICING=anthropic/claude-sonnet-4-6
OPENROUTER_MODEL_BOOKING=deepseek/deepseek-v4-proTwo separate settings, because the two jobs have different requirements. The extraction job produces long structured output against a complex schema. The quote helper produces three sentences. There is no reason to pay frontier prices for the second one.
Changing either is an edit to an environment file and a page reload. No redeploy, no code change, no SDK swap. When a new model appears, evaluating it on real data is a two-minute experiment, and that has changed how often I bother to check.
Two practical notes about the API itself. The request is not streamed, because there is nothing to show progressively; the caller wants a complete object or an error. And the timeout is 600 seconds for extraction, which sounds absurd until you watch a reasoning model take four minutes on a property with fifteen seasons and three holiday packages.
Long calls and session locks#
An unglamorous bug worth mentioning, because it will bite anyone doing this in PHP.
The extraction endpoint sits behind session-based authentication. PHP’s session_start() holds an exclusive lock on the session file for the whole request. So a four-minute AI call meant that every other request from the same logged-in user, in any other tab, queued behind it. The dashboard appeared to hang, and each waiting request pinned a worker of its own.
The fix is one line, placed after the auth and CSRF checks and before the AI call:
if (session_status() === PHP_SESSION_ACTIVE) session_write_close();What goes wrong with the JSON#
“Output only JSON” is an instruction, not a type system. Over a few thousand runs the failures sort into four groups.
| Failure | What arrives | Fix |
|---|---|---|
| Markdown fences | The object inside a code fence, maybe with a sentence before it | Take the first { to the last } |
| Typographic punctuation | Curly quotes, trailing commas | Normalise before parsing |
| Reasoning tags | <thinking> blocks ahead of the answer |
Strip them before parsing |
| Truncation | Valid JSON that stops mid-object | Raise the limit, then close open containers |
Markdown fences. The single most common one. The object arrives wrapped in a fenced code block, sometimes with a language tag, sometimes with a friendly sentence before it. Extracting from the first { to the last } handles nearly all of it.
Typographic punctuation. Some models emit curly quotes inside JSON, particularly when the surrounding text is not English. A three-line normaliser deals with it, along with trailing commas before a closing brace:
function cleanupAiJsonString(string $s): string {
$s = str_replace(["\xE2\x80\x9C", "\xE2\x80\x9D", "\xE2\x80\x9E"], '"', $s);
$s = str_replace(["\xE2\x80\x98", "\xE2\x80\x99", "\xE2\x80\x9A"], "'", $s);
$s = preg_replace('/,(\s*[\]}])/', '$1', $s); // trailing commas
return $s;
}Reasoning tags. Models with visible chain-of-thought sometimes wrap it in <thinking> or <reasoning> blocks and put the answer after. Stripping those blocks before parsing is a two-line regex and turns a failed run into a good one.
Truncation. The interesting one. A long property runs out of output tokens mid-object, and you get perfectly valid JSON for the first eighty percent followed by nothing.
The first response was to raise the limit, from 8192 to 24000 tokens. That was correct and mostly sufficient, and it costs nothing extra because you are billed for tokens generated, not tokens allowed. But “mostly” is not “always”, so there is also a repair step, and its design matters more than it looks.
The naive repair is to trim back to the last point where braces balance. That is wrong here, and wrong in an expensive way: a truncated object never balances at depth zero, so trimming back leaves you with an inner array, and the top-level property block that carries capacity and tax settings is gone. The importer then writes a property record full of defaults.
So instead the repair walks the string, tracks the open containers, and appends the closing brackets the text still owes:
function closeOpenJsonContainers(string $json): ?string {
$stack = []; $inString = false; $escape = false;
for ($i = 0, $n = strlen($json); $i < $n; $i++) {
$ch = $json[$i];
if ($escape) { $escape = false; continue; }
if ($inString) {
if ($ch === '\\') $escape = true;
elseif ($ch === '"') $inString = false;
continue;
}
if ($ch === '"') { $inString = true; continue; }
if ($ch === '{') $stack[] = '}';
elseif ($ch === '[') $stack[] = ']';
elseif ($ch === '}' || $ch === ']') array_pop($stack);
}
if ($inString || empty($stack)) return null;
return rtrim(rtrim($json), ',') . implode('', array_reverse($stack));
}If the result still does not parse, because the cut landed on a dangling key or a half-written value, the loop drops the last element and tries again.
The whole repair path is maybe 120 lines and it is the least interesting code in the feature. It is also the reason the feature can be trusted, because it converts a class of silent corruption into either a correct import or a visible failure.
Validate everything the model returns#
After parsing comes a normalisation pass that treats the JSON as hostile input, which it is: it comes from a stochastic process over text an owner typed.
Rows are dropped when they cannot possibly work: a season without a valid date range or with a base price of zero, a discount with a type not in the enum, a rule whose dates are reversed. Values are clamped: percentages above 100, negative fees, weekday lists containing an 8. Cross-references are resolved and cleared when they dangle, so a holiday linked to a season that does not exist loses the link instead of silently linking to nothing.
Some errors get repaired rather than dropped, where the intent is unmistakable. A discount with start_date after end_date is nearly always a transposition, and dropping it turns a date-limited promotion into an always-on one, which is worse than fixing it:
if (!empty($d['start_date']) && !empty($d['end_date'])
&& $d['end_date'] < $d['start_date']) {
[$d['start_date'], $d['end_date']] = [$d['end_date'], $d['start_date']];
}Then the last and most important gate: nothing is written until a person looks at it. The parsed result is rendered as a table of seasons, holidays, discounts and rules, in the operator’s language, next to the original text. They approve the import or they do not. Rows an operator has previously locked by hand are never overwritten, so the import can only add around them.
One call per year#
Price lists cover several years, often in one page with headings. Sending all of it in one request produced two reliable problems: dates from the wrong year, and truncation.
Splitting the text on headings that contain a four-digit year and sending one request per section fixed both. Each call sees one year, is told which year it is looking at, and returns a smaller object. Results are merged, then deduplicated on natural keys, because the same holiday sometimes appears in two sections.
The interesting part is that this made the results better rather than merely smaller. A prompt that says “the input is one year, use this year for every date” removes an entire category of reasoning the model was doing badly. It is the same lesson as splitting a large function: less context, fewer ways to be wrong.
Which model#
I have run this task across several tiers, and the differences show up in specific places rather than as a general quality score.
| Tier | Strength | Watch out for | Use it for extraction? |
|---|---|---|---|
| Frontier commercial | Complex schemas right first time | Price, latency | Yes, when quality is everything |
| Mid-tier commercial | The common case, cleanly | Awkward edge rules | Yes, for simple properties |
| Large open-weight | Cheap, fast, good on clear text | Needs the full safety net | Yes, the current choice |
| Reasoning-heavy | Little, for this task | Token count, latency, tags | No |
| Small and cheap | Short summaries | Plausible wrong values | No |
Frontier commercial models get complex schemas right the first time. Nested conditional rules, the two conflicting end-date conventions, holidays that inherit a season’s price: they follow all of it, and the repair path rarely fires. They are also the most expensive per token and often the slowest, particularly if they reason before answering.
Mid-tier commercial models handle the common case perfectly well. Simple properties with three seasons and a cleaning fee come back clean. The failures cluster around the awkward parts: the weekday and weekend split into two rows, the priority field on nested seasons, and the exact substring the engine matches on. For many use cases this is the right trade, at roughly half the price or less.
Large open-weight models have become genuinely competitive for this task, which is the main reason the config currently points at one. They are an order of magnitude cheaper, they are fast, and on a clear price list the output is indistinguishable from the expensive option. They need more of the safety net: markdown fences more often, truncation more often, and instructions taken more literally, which cuts both ways. An instruction like “use null when unknown” is followed exactly; an instruction that relies on inference is followed less reliably.
Reasoning-heavy models were the surprise, and not a pleasant one. On a task like this the reasoning does not help much, because the work is transcription against a fixed schema rather than deduction. What it does is multiply the output tokens by several times and the latency by more. It also produced the <thinking> tag mess that needed handling. My conclusion, held loosely, is that extraction wants a model that follows instructions precisely, not one that thinks hard.
Small and cheap models are not usable for the extraction, and the failure mode is instructive: the JSON is syntactically fine, the field names are right, and the values are wrong in ways that look plausible. That is the worst possible outcome, far worse than a parse error. They are perfectly good for the short quote-summary job.
Two practical points. Prompts are not fully portable between families, and a prompt tuned for one model will lose a few percent on another, though a very explicit prompt with worked examples travels better than a terse one. And OpenRouter lets you evaluate all of this by changing a string, which is the only reason I have opinions about five models instead of one.
What it costs#
The unit economics are the part people usually get wrong in both directions.
Here is the arithmetic for one property. The prompt is around 5,000 tokens, most of it the fixed schema and rules. The property’s price text adds 500 to 2,000 tokens. Output is 1,000 to 4,000 tokens of JSON. A property covering three years is three of those calls, so call it 25,000 input tokens and 9,000 output tokens for a complete property.
List prices, per million tokens, at the time of writing, and what they come to for that one property:
| Tier | Input | Output | One property |
|---|---|---|---|
| Frontier commercial | $5 and up | $25 and up | about $0.35 and up |
| Mid-tier commercial | around $3 | around $15 | about $0.20 |
| Small commercial | around $1 | around $5 | about $0.07 |
| Large open-weight | cents to low dollars | cents to low dollars | about $0.01 |
Prices move, providers differ, and OpenRouter shows the current number on each model’s page, so treat that table as a shape rather than a quote. Leaving out the small models, which are not usable for this job anyway, a full property extraction runs from roughly a cent on an open-weight model, through about twenty cents mid-tier, to thirty-five cents or more at the top.
A hundred properties, once a year, is somewhere between one dollar and thirty-five. The manual alternative is twenty minutes of a person’s attention per property, which is thirty-three hours.
That comparison is so lopsided that the cost question stops being about money. The real costs are elsewhere:
Latency. Four minutes for a multi-year property is a genuine UX problem, and the honest fix is not a faster model. It is a job queue, a progress indicator, and not blocking anything else while it runs.
Review time. A person still reads every result. If that takes five minutes instead of twenty, the feature has paid for itself many times over. If a cheap model’s error rate pushes it back toward fifteen, the saving on tokens was an illusion. This is the number to optimise, and it is why I would rather pay four times more per token than lose the operator’s trust.
The engineering around it. Repair, validation, per-year splitting, the review UI, the locked-row logic. Call it 1,500 lines. That is the actual price of the feature, and it is a one-time cost that no choice of model reduces.
What I would tell someone building this#
- Write the prompt like a specification, not like a request. Every ambiguity you leave is a decision the model makes freshly on each run.
- Never let the model do the arithmetic that produces a number a customer sees. Extraction into a deterministic engine gives you the model’s flexibility with the engine’s reliability.
- Assume the JSON is broken and handle it in layers: extract the object from surrounding text, normalise the punctuation, repair truncation, validate every field, then show a human a table. Each layer converts a silent failure into either a correct result or a visible one.
- Use OpenRouter and keep the model name in configuration. One OpenAI-compatible endpoint lets you try many models on real data, and switching or migrating to a new one is a config change, not a code change.
- Keep a person on the last step for anything that touches money. Not because the model is bad, but because a review that takes a minute makes the difference between a feature people use and a feature people quietly stop trusting.