Jev and where a System One model fits in document processing → Unstract.com

· Unstract.com → ·

58 min read Original article ↗

Note: Pricing in this article is as of September 2026. TypeSafe lists Jev at $0.042 per million input tokens with output tokens free, and every cost figure below is worked out from that. Check the TypeSafe docs for current prices.

What is Jev?

TypeSafe AI released Jev in September 2026. They call it a System One model, a new class of model meant to be called by code, not chatted with by people. The name comes from Kahneman’s System 1: the fast, intuitive kind of thinking that doesn’t stop to deliberate.

LLMs are trained with Reinforcement Learning from Human Feedback (RLHF). That is what makes them good at following instructions from humans. Jev is trained with something TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD).

The goal is not to follow instructions well. The goal is to give honest probabilities. When we say calibrated, it means when Jev says 0.9, it should be right about 90% of the time across many such answers. 

What Jev cannot do

Jev cannot generate free form text like how LLMs do. It cannot chat with you. It cannot reply in plain English. Most importantly from this article’s perspective of document processing, it cannot extract a value from a document.

You cannot ask it for the invoice number and it has no way to write "INV-2026-0428" back to you. It cannot summarise. It cannot suggest a better alternative. It cannot read images or PDFs either. It takes text only for now.

What Jev can do

Given a context (TypeSafe calls it the “state”), like the text of an invoice, it can answer three types of questions:

1. Choice / Pick one option from a list

Up to 255 options.

“What is this invoice billing for?”

  • [Goods, Services, Mixed, Unclear]
    ↓
    Jev returns a probability (0.0–1.0) for each option, the winning option, and a confidence number that says how peaked the distribution is.

2. Score / Rate the state against an ordered rubric

2 to 10 levels.

“How specific is the description of the items billed?”

  • 0: No meaningful description.
  • 1: General service named, but scope unclear.
  • 2: Specific activities and scope described.
    ↓
    Jev returns a probability for each level, plus a single score that can land between levels.

A 1.4 means it is leaning towards 2 but not sure.

3. Noul / Is this statement true?

“Does the invoice explicitly state when payment is due?”

  • [no list needed]
    ↓
    Jev returns one number: the probability that the statement is true. Nothing else.

You can put many questions of all three kinds into one call. TypeSafe says a response takes 70–500 ms and costs $0.042 per million input tokens, with output free as of Sept 2026. Even asking a question about every field on every document becomes affordable.

Why this is useful

Fundamentally, Jev only answers closed questions. Every possible answer has to be part of your ask. An LLM can answer an open question: “what is the invoice number?” but Jev cannot because it would have to write the answer out, and it doesn’t write.

But if you already have a candidate, from an LLM, a regex, or a lookup in your vendor master for the invoice example, Jev can tell you how likely it is that the document supports/has it. Or it can pick the right one from a list. Or tell you which schema/class this document belongs to.

TypeSafe also says Jev “can’t hallucinate”. It only means the answer will always be one of your options. It can still confidently pick the wrong one.

This is a huge limitation. But document pipelines have requirements of exactly this kind of closed questions, and today they are answered by an LLM that is slower, far more expensive, and grades its own work. That is what this article is about. We will go through six use cases where it is useful.


Six use cases to start with

Jev shows up in three places in a document pipeline: after the LLM has produced its answer, before the LLM runs, and instead of the LLM. We start with the one that matters most, then go through the pipeline from the front.

Every example follows the same format: a sample document, the text LLMWhisperer produced from it, the question we asked Jev, what Jev returned, what that buys you and what it doesn’t. All documents are synthetic, built for our demo kits, with the traps planted on purpose.

After extraction

1. Verify every field, escalate the doubtful ones.

The LLM has extracted a record and rated its own confidence. One Jev call asks, for every field, whether the document actually supports that value. The fields it doubts go for re-extraction or to a person. If you are going to read only one section, read this one.

Before extraction

2. Triage: what is this document?

An accounts payable inbox gets invoices, but also pro formas, statements and credit notes. A pro forma that reaches the invoice prompt extracts perfectly and gets paid. One Choice question decides which prompt runs, before anything is spent.

3. Packet pages: which pages matter, and do they belong here?

A loan packet arrives as one 17-page fax with six kinds of documents in it, and one page belongs to someone else. A Choice and a Noul per page split the packet, route each piece to its own schema, and catch the stray page.

Instead of extraction

4. Closed-set fields: skip the LLM.

Many fields in a schema are already yes/no or one-of-N. Citizenship, loan purpose, occupancy, the declarations on a loan application. The form prints the options, so Jev answers these directly from the text, with a probability you can act on, and the LLM only handles the fields that need a value written out.

5. Candidate selection: several values, one field.

A policy change form prints three dates. Code finds them all, Jev points at the right one, and code copies it verbatim. Jev never writes the value, so it cannot transpose a digit.

6. Mapping to a master list.

A carrier’s loss run says SUBRO PEND. Your taxonomy has five statuses. Jev maps each row to your list, with unmapped as a legitimate answer instead of a forced fit.

Stop wrestling with complex documents — See how LLMWhisperer parses them


The layout-preserved text that feeds our agent pipeline comes from LLMWhisperer. Try it on your own messy PDFs — bank statements, rent rolls, multi-header tables, cross-page rows — No signup required.

Try LLMWhisperer for free on the Playground. No signup required.


Use case 1: Verify every field, escalate the doubtful ones

The problem

The LLM has read the document and returned a JSON record. Some of the fields are right. Some are wrong. You don’t know which, and neither does the LLM. It rates its own confidence, and that number is not reliable. Your options today are not great. A second LLM as a challenger costs about as much as the extraction did. A person can’t read every field of every document. So most pipelines sample, or threshold on a confidence number that means very little.

Jev can ask “does the document has this value?” for every field, in one call, for a fraction of a cent.

The document

A commercial auto loss run from a carrier. Two pages with five claims. This is the kind of document an insurance broker receives dozens of a day and needs to turn into rows.

Look at the last claim (columns trimmed to fit):

Claim No    Claimant / Description       Stat      Paid  O/S Reserve  Incurred
SBI-556810  Devon Mackey — Cargo theft   CLSD  9,400.00         0.00  8,200.00

Paid is 9,400. Reserve is 0. Incurred is 8,200.
That doesn’t add up, and it isn’t supposed to: the carrier’s own report has a transposition in it. There is no recovery column anywhere on this document.

The last line says “Totals not provided on this report.” An LLM that knows the identity incurred = paid + reserve − recovery has an easy way out. It can invent a recovery of 1,200 to make the numbers balance. Nothing on the page says 1,200. It is a number that makes the arithmetic work, and that is exactly the kind of error that gets past every check you already have.

LLMWhisperer output

LLMWhisperer v2, layout-preserving output, native_text mode (the PDF is born-digital; a scan would go through form or high_quality mode). 2,812 characters for the two pages, about 700 tokens. 

Sentinel Bay Insurance Group                                                                                                                                                                                         Page 1 
NAIC 99358  ·  Claims Services  ·  CA                                                                                                                                                             Valued as of 07/24/2026 

COMMERCIAL AUTOMOBILE — CLAIMS DETAIL 

Named Insured:   Cascade Millwork LLC 
DBA:             Cascade Cabinetry 
FEIN:            00-9100041 
Address:         17585 SW Alderfen Loop, Tualatin, OR 97062 
Policy Number:   AU-7781-0044 
Policy Period:   08/11/2025 to 08/10/2026 
Coverage:        Commercial Auto 

Claim No               DOL             Rept             Claimant / Description                                                   Cause                 Stat            Paid       O/S Reserve                Incurred 

SBI-556201             10/05/2025      10/06/2025       Vernon Pike — Rear-end collision, box truck; third-party bodily injury   BI-AUTO               OPEN        6,000.00           5,900.00              11,900.00 
SBI-556388             11/29/2025      11/30/2025       — — Backing collision at loading dock, own unit                          COLL                  CLSD        3,300.00               0.00                3,300.00 
<<<

Sentinel Bay Insurance Group                                                                                                                                                                                         Page 2 
NAIC 99358  ·  Claims Services  ·  CA                                                                                                                                                             Valued as of 07/24/2026 

COMMERCIAL AUTOMOBILE — CLAIMS DETAIL 

Claim No               DOL             Rept             Claimant / Description                                                   Cause                 Stat            Paid       O/S Reserve                Incurred 

SBI-556502             02/02/2026      02/04/2026       Ana Souza — Parking lot sideswipe, third-party property damage           PD-AUTO               SUBRO PEND 2,900.00                0.00                2,900.00 
SBI-556677             04/13/2026      04/14/2026       — — Windshield replacement, comprehensive                                GLASS                 CLSD          900.00               0.00                  900.00 
SBI-556810             06/02/2026      06/04/2026       Devon Mackey — Cargo theft from parked unit                              THEFT                 CLSD        9,400.00               0.00                8,200.00 

End of detail. Totals not provided on this report. 
<<<

Column alignment matters here. Jev has to be able to tell that 9,400.00 sits under Paid and not under Incurred, and it can only do that if the text keeps the columns where the page had them. This text is the “state” in every Jev call below.

The extraction record

This is the record an extraction step hands to the rest of the pipeline: one object per claim, in the shape our loss-run extraction prompt asks for. For this example we wrote it by hand and planted one wrong value, the recovery of 1,200 on SBI-556810. Everything else is copied from the report. Here is that claim:

{
  "claim_number": "SBI-556810",
  "date_of_loss": "2026-06-02",
  "date_reported": "2026-06-04",
  "claimant_name": "Devon Mackey",
  "loss_description": "Cargo theft from parked unit",
  "cause_raw": "THEFT",
  "status_raw": "CLSD",
  "paid_indemnity": 9400.00,
  "paid_alae": null,
  "reserve_indemnity": 0.00,
  "reserve_alae": null,
  "gross_incurred": 8200.00,
  "recovery": 1200.00,
  "source_page": 2
}

Two things to notice. recovery is 1,200, and the document never says so. And paid_alae  (Allocated Loss Adjustment Expenses) is null, which is correct: this carrier prints one Paid column, not a split.

A field the document doesn’t have should stay empty. Both need checking, and they are different checks: one asks “is this value on the page?”, the other asks “is there something on the page that should be here?”

The question to Jev

Ten questions about this one claim, in one call. In production you would ask about every field of every claim in the same call (we did that too: 60 questions, one second), but ten is enough to print the whole exchange here.

Every question is generated by code from the record. For a field with a value, ask whether the report prints it. For an empty field, ask whether the report prints something there.

from typesafe_sdk import Noul, TypeSafeClient

def money(v):   # write it the way the report prints it; Jev reads literally
    return f"{float(v):,.2f}"

def date(v):    # 2026-06-02 -> 06/02/2026
    return f"{v[5:7]}/{v[8:10]}/{v[0:4]}"

cn = claim["claim_number"]
questions = {
  "date_of_loss":      Noul(instructions=f"The report prints {date(claim['date_of_loss'])} "
                                         f"as the DOL (date of loss) for claim {cn}"),
  "claimant_name":     Noul(instructions=f"The report prints {claim['claimant_name']} "
                                         f"as the claimant for claim {cn}"),
  "loss_description":  Noul(instructions=f"The report prints {claim['loss_description']} "
                                         f"as the description for claim {cn}"),
  "cause_raw":         Noul(instructions=f"The report prints {claim['cause_raw']} "
                                         f"as the Cause code for claim {cn}"),
  "status_raw":        Noul(instructions=f"The report prints {claim['status_raw']} "
                                         f"as the Stat (status) for claim {cn}"),
  "paid_indemnity":    Noul(instructions=f"The report prints {money(claim['paid_indemnity'])} "
                                         f"as Paid for claim {cn}"),
  "reserve_indemnity": Noul(instructions=f"The report prints {money(claim['reserve_indemnity'])} "
                                         f"as O/S Reserve for claim {cn}"),
  "gross_incurred":    Noul(instructions=f"The report prints {money(claim['gross_incurred'])} "
                                         f"as Incurred for claim {cn}"),
  # an empty field: is there something on the page that should have been extracted?
  "paid_alae_absent":  Noul(instructions=f"The report has an ALAE (expense) paid column, "
                                         f"and the row for claim {cn} has a figure in it"),
  # the planted value: is 1,200.00 actually printed as a recovery?
  "recovery":          Noul(instructions=f"The report has a Recovery column, and the row "
                                         f"for claim {cn} prints {money(claim['recovery'])} in it"),
}

response = TypeSafeClient().system_one(
    model="jev-1.13.0", state=whisperer_text, questions=questions)

The request that was actually sent to Jev by the SDK:

{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text shown above, 2,812 characters>",
  "questions": {
    "date_of_loss": {
      "type": "noul",
      "instructions": "The report prints 06/02/2026 as the DOL (date of loss) for claim SBI-556810"
    },
    "claimant_name": {
      "type": "noul",
      "instructions": "The report prints Devon Mackey as the claimant for claim SBI-556810"
    },
    "loss_description": {
      "type": "noul",
      "instructions": "The report prints Cargo theft from parked unit as the description for claim SBI-556810"
    },
    "cause_raw": {
      "type": "noul",
      "instructions": "The report prints THEFT as the Cause code for claim SBI-556810"
    },
    "status_raw": {
      "type": "noul",
      "instructions": "The report prints CLSD as the Stat (status) for claim SBI-556810"
    },
    "paid_indemnity": {
      "type": "noul",
      "instructions": "The report prints 9,400.00 as Paid for claim SBI-556810"
    },
    "reserve_indemnity": {
      "type": "noul",
      "instructions": "The report prints 0.00 as O/S Reserve for claim SBI-556810"
    },
    "gross_incurred": {
      "type": "noul",
      "instructions": "The report prints 8,200.00 as Incurred for claim SBI-556810"
    },
    "paid_alae_absent": {
      "type": "noul",
      "instructions": "The report has an ALAE (expense) paid column, and the row for claim SBI-556810 has a figure in it"
    },
    "recovery": {
      "type": "noul",
      "instructions": "The report has a Recovery column, and the row for claim SBI-556810 prints 1,200.00 in it"
    }
  }
}

What Jev returns

0.92 seconds. 1,281 input tokens, 188 output tokens. The response, verbatim:

{
  "model": "jev-1.13.0",
  "usage": {
    "input_tokens": 1281,
    "output_tokens": 188
  },
  "answers": {
    "date_of_loss":      {"type": "noul", "noul": 0.99},
    "claimant_name":     {"type": "noul", "noul": 0.93},
    "loss_description":  {"type": "noul", "noul": 0.98},
    "cause_raw":         {"type": "noul", "noul": 0.99},
    "status_raw":        {"type": "noul", "noul": 0.99},
    "paid_indemnity":    {"type": "noul", "noul": 0.99},
    "reserve_indemnity": {"type": "noul", "noul": 0.99},
    "gross_incurred":    {"type": "noul", "noul": 0.99},
    "paid_alae_absent":  {"type": "noul", "noul": 0.03},
    "recovery":          {"type": "noul", "noul": 0.02}
  }
}

Read against the record:

Field Extracted value Question Jev Action
date_of_loss 2026-06-02 P(printed as DOL) 0.99 accept
claimant_name Devon Mackey P(printed as claimant) 0.93 accept
loss_description Cargo theft from parked unit P(printed as description) 0.98 accept
cause_raw THEFT P(printed as Cause) 0.99 accept
status_raw CLSD P(printed as Stat) 0.99 accept
paid_indemnity 9,400.00 P(printed as Paid) 0.99 accept
reserve_indemnity 0.00 P(printed as O/S Reserve) 0.99 accept
gross_incurred 8,200.00 P(printed as Incurred) 0.99 accept
paid_alae null P(ALAE paid column has a figure) 0.03 accept null
recovery 1,200.00 P(Recovery column prints 1,200.00) 0.02 review

Out of the ten questions, only one came back with a low probability and that is the one we planted. Keep in mind that 1,200 is a perfectly reasonable value for a recovery field. It is a number, it is in the right field and it makes the row add up. Any check that looks at the record alone will pass it. Jev was asked whether there is a Recovery column with 1,200.00 in it for this claim. There isn’t, so it returned 0.02.

What this buys you

You can now check every field of every document instead of sampling. The ten questions took under a second and 1,281 input tokens. At TypeSafe’s published price that is about $0.00005.

The full version with all five claims and 60 questions took 1.2 seconds and cost about $0.00012. The other thing you get is a second opinion from a different model. The extractor is not grading its own work anymore.

We also ran the actual extraction with LLMs on this document, 17 times with three different models, to see how often this recovery gets invented. It happened in 6 of the 17 runs.

The confidence the LLM gave that claim row ran from 0.62 to 0.85 on the runs where it invented the recovery and from 0.70 to 0.95 on the runs where it did not. No threshold separates the two. So the LLM’s own confidence would not have caught it. Jev gave the invented recovery 0.02 or 0.03 in all 6 runs.

What it doesn’t buy you

Jev reads what LLMWhisperer gave it, not the actual document. If the OCR reads a 3 as an 8, the extraction will have the 8, Jev will confirm the 8 and everything passes. For that you need LLMWhisperer’s per word confidence metadata, which tells you which words the OCR was not sure about. That is a different check and we will cover it in a separate article.

How to turn complex document tables into usable data with AI

Catch the recorded webinar to dive into Unstract’s advanced table extraction capabilities. We walk through the All Table Extractor API—a ready-to-use, semantic + layout-aware solution that ensures precision, consistency, and context, even across the most complex document tables.


Use case 2: Triage: what is this document?

The problem

Everything that lands in an accounts payable inbox is going to cost you something to process. Before you run an extraction prompt on it, you want to know what it is. Is it an invoice, a statement, a credit note, a delivery note? Which prompt should run on it, or should one run at all?

The document that hurts you here is not the stray statement or delivery note. Those extract badly and fail somewhere downstream. The one that hurts is the document that extracts perfectly into the wrong record.

The documents

Two documents from the same supplier, Alder Creek Electric. Same letterhead, same layout, same footer, same payment terms. Both quote a real, open purchase order and both have line items whose codes, quantities and prices match that order exactly.

One is an invoice and the other is a pro forma, which is a priced offer, not a demand for payment. If the pro forma reaches the invoice prompt, it extracts cleanly, the PO match passes and a bill gets posted against a document that was never an invoice. When the real invoice arrives you have a duplicate at best and a double payment at worst.

LLMWhisperer output

LLMWhisperer v2, layout preserving output, native_text mode since both PDFs are born digital. About 1,800 characters each, roughly 450 tokens.

Alder Creek Electric, Inc.                                                                       INVOICE 

5850 NE Ivyfield Boulevard 
Portland, OR 97218                                                           Invoice no.                  ACE-2291 
(503) 555-0142  ·  [email protected]                           Invoice date                09/04/2026 
Federal Tax ID  00-2185540 
                                                                             Your order                    PO-4471 

                                                                             Terms                           Net 30 

                                                                             Currency                          USD 

BILL TO 
Talbot Facilities Group, Inc. 
Accounts Payable 
1695 SW Pikewell Avenue, Suite 400 
Portland, OR 97204 

ITEM                 DESCRIPTION                                      QTY             UNIT PRICE            AMOUNT 

EL-COND-20           20mm steel conduit, 3m length                     40                $19.60             $784.00 

EL-JBOX-4W           4-way junction box, IP65                          25                $42.40           $1,060.00 

                                                                                        Subtotal          $1,844.00 

                                                                                       Sales tax              $0.00 

                                                                                    TOTAL DUE             $1,844.00 

Payment due 30 days from invoice date. Please quote the invoice number with your 
remittance. Questions: (503) 555-0142 or [email protected]. 

Cascade Document Services  ·  batch TG-0909-A  ·  item 0001  ·  page 1 of 1 
<<<
Alder Creek Electric, Inc.                                            PRO FORMA INVOICE 

5850 NE Ivyfield Boulevard 
Portland, OR 97218                                                           Invoice no.                ACE-PF-118 
(503) 555-0142  ·  [email protected]                           Invoice date                09/01/2026 
Federal Tax ID  00-2185540 
                                                                             Your order                    PO-4544 

                                                                             Terms                           Net 30 

                                                                             Currency                          USD 

BILL TO 
Talbot Facilities Group, Inc. 
Accounts Payable 
1695 SW Pikewell Avenue, Suite 400 
Portland, OR 97204 

ITEM                 DESCRIPTION                                      QTY             UNIT PRICE            AMOUNT 

EL-TRAY-150          150mm cable tray, 3m length                       60                $31.80           $1,908.00 

                                                                                        Subtotal          $1,908.00 

                                                                                       Sales tax              $0.00 

                                                                                    TOTAL DUE             $1,908.00 

Payment due 30 days from invoice date. Please quote the invoice number with your 
remittance. Questions: (503) 555-0142 or [email protected]. 

Cascade Document Services  ·  batch TG-0909-A  ·  item 0008  ·  page 1 of 1 
<<<

Read them side by side. Both say “Invoice no.” Both say “TOTAL DUE”. Both say “Payment due 30 days from invoice date”.

The pro forma gives itself away in exactly two places, the title and the PF in its number. Everything structural about the two documents is the same. A rule that looks for the words “total due” or for a PO number passes both.

The question to Jev

One call per document. A Choice for the kind of document, and three Nouls that the routing code needs anyway. The Nouls ride in the same call at no extra latency, so there is no reason to ask them separately.

from typesafe_sdk import Choice, Noul, TypeSafeClient

QUESTIONS = {
    "kind": Choice(
        instructions="What kind of document is this?",
        criteria={
            "invoice":       "A demand for payment for goods or services already supplied",
            "pro_forma":     "A priced offer, quotation or estimate issued before supply. "
                             "It may be laid out like an invoice but it is not a demand for payment",
            "credit_note":   "A supplier reducing an amount it previously invoiced",
            "statement":     "A summary of an account's open items, usually listing several invoices",
            "delivery_note": "A record of goods delivered, without a payment demand",
            "other":         "None of the above",
        }),
    "quotes_po":       Noul(instructions="The document quotes a customer purchase order number"),
    "has_line_items":  Noul(instructions="The document contains a table of itemised charges "
                                         "with quantities and prices"),
    "addressed_to_us": Noul(instructions="The document is addressed to Talbot Facilities Group"),
}

for doc in inbox:
    response = TypeSafeClient().system_one(
        model="jev-1.13.0", state=doc.whisperer_text, questions=QUESTIONS)

The request, exactly as sent for the invoice. The pro forma’s request is identical except for the state:

{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text for 01-alder-creek-ace-2291.pdf, 1,811 characters>",
  "questions": {
    "kind": {
      "type": "choice",
      "instructions": "What kind of document is this?",
      "criteria": {
        "invoice": "A demand for payment for goods or services already supplied",
        "pro_forma": "A priced offer, quotation or estimate issued before supply. It may be laid out like an invoice but it is not a demand for payment",
        "credit_note": "A supplier reducing an amount it previously invoiced",
        "statement": "A summary of an account's open items, usually listing several invoices",
        "delivery_note": "A record of goods delivered, without a payment demand",
        "other": "None of the above"
      }
    },
    "quotes_po": {
      "type": "noul",
      "instructions": "The document quotes a customer purchase order number"
    },
    "has_line_items": {
      "type": "noul",
      "instructions": "The document contains a table of itemised charges with quantities and prices"
    },
    "addressed_to_us": {
      "type": "noul",
      "instructions": "The document is addressed to Talbot Facilities Group"
    }
  }
}

The option descriptions matter. Jev reads literally, so “may be laid out like an invoice but it is not a demand for payment” is there to stop the layout from deciding the answer.

What Jev returns

Invoice: 0.95 seconds, 865 input tokens. Pro forma: 0.89 seconds, 836 input tokens. Both responses, verbatim:

01-alder-creek-ace-2291 (the invoice)

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 865, "output_tokens": 118},
  "answers": {
    "kind": {
      "type": "choice",
      "choice": "invoice",
      "confidence": 1.0,
      "probabilities": {"invoice": 1.0, "pro_forma": 0.0, "credit_note": 0.0,
                        "statement": 0.0, "delivery_note": 0.0, "other": 0.0}
    },
    "quotes_po":       {"type": "noul", "noul": 0.99},
    "has_line_items":  {"type": "noul", "noul": 0.99},
    "addressed_to_us": {"type": "noul", "noul": 0.99}
  }
}

08-alder-creek-ace-pf-118 (the pro forma)

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 836, "output_tokens": 120},
  "answers": {
    "kind": {
      "type": "choice",
      "choice": "pro_forma",
      "confidence": 0.93,
      "probabilities": {"pro_forma": 0.94, "invoice": 0.06, "credit_note": 0.0,
                        "statement": 0.0, "delivery_note": 0.0, "other": 0.0}
    },
    "quotes_po":       {"type": "noul", "noul": 0.99},
    "has_line_items":  {"type": "noul", "noul": 0.99},
    "addressed_to_us": {"type": "noul", "noul": 0.99}
  }
}
Question 01 invoice 08 pro forma
kind invoice 1.00 pro_forma 0.94 (invoice 0.06)
confidence 1.00 0.93
quotes_po 0.99 0.99
has_line_items 0.99 0.99
addressed_to_us 0.99 0.99
route extract:invoice log_and_skip:pro_forma

Three of the four answers are the same for both documents. That is expected, the two documents have the same shape.

The only answer that differs is the kind, and that is the one that decides what happens next. We ran the same call on all 8 documents in this inbox, which has six more invoices, five from other suppliers and a second copy of the Alder Creek one, one of them a scan that went through LLMWhisperer’s OCR mode.

All 7 invoices came back invoice at 1.00 and the pro forma came back pro forma at 0.93.

What this buys you

An LLM can classify documents too. The reason to use Jev for this is where the step sits. It runs on every single document, before you have decided to spend anything on it. At TypeSafe’s published price the call above is about $0.00004 per document and it takes under a second.

The Nouls are the other reason. Every yes/no your routing code needs –  does it quote a PO?, does it have line items?, is it addressed to us?, is it in USD?, is there a credit mentioned?, can be asked in the same call. That is a dozen branches in your pipeline without a dozen model calls.

What it doesn’t buy you

Jev classifies on what the text says, and here the title is the only thing it is going by. We checked this. We blanked out the words PRO FORMA INVOICE from the pro forma’s text, left the ACE-PF-118 number in, and asked the same question again.

Jev said invoice, at 1.00. The PF in the number did nothing. So a supplier who sends a quotation and calls it an invoice will get through Jev, and through an LLM for that matter. Nothing in the text separates it from an invoice anymore. That is what the three way match downstream is for, and it is why use case 1 exists.

One more thing worth knowing. When the tell was removed, Jev did not hedge. It did not come back with 0.5 invoice and 0.4 pro forma. It came back with invoice at 1.00 and confidence 1.00. Calibrated means the probabilities are honest about what is in the text. It does not mean the model knows what is missing from the text.

Also, these are single page documents. A 40 page bundle where page 12 is the invoice is a different problem, which is use case 3.

Turn your complex PDF tables into structured data with Unstract


Unstract uses LLMs to extract clean, structured JSON from any document — PDFs, scans, images, tables of any layout. Define what you want using natural language, deploy as an API or ETL pipeline, and get data your systems can actually use.

Try Unstract for free on the Playground. No signup required.


Use case 3: Packet pages: which pages matter, and do they belong here?

The problem

A loan packet does not arrive as nine documents. It arrives as one PDF, 17 pages, faxed from a broker’s office. Somewhere in it there is a loan application, a driver license, pay stubs, W-2s, bank statements and a purchase agreement. Each one needs a different extraction schema, and one of them may not belong to this borrower at all.

You cannot hand the whole packet to one extraction prompt. It is too big, the schemas are different, and Jev’s own  documentation says accuracy drops when the state contains material that has nothing to do with the question. So before any extraction runs, you need to know what each page is, where each document starts and stops, and whose documents these are.

The document

quaye-package.pdf, 17 pages. A fax rescan, grey, slightly skewed, with a fax banner across the top of every page. It contains nine documents:

The Mix:

Pages Document
1–5 Uniform Residential Loan Application (Form 1003)
6 Oregon driver license
7 Pay stub, Sandy River Dental, for Denise Achterberg
8 Pay stub, Westside Transit Services, for Bartholomew Quaye
9 W-2, 2025
10 W-2, 2024
11–12 Bank statement, August
13–14 Bank statement, July
15–17 Residential purchase agreement

Page 7 is the trap. The broker’s office scanned two clients’ documents into one packet. Page 7 is a pay stub for Denise Achterberg. Page 8 is a pay stub for Bartholomew Quaye, the borrower. Same kind of document, next to each other, and one of them has no business being here. A required document that names nobody on the loan means the whole packet goes back for review.

LLMWhisperer output

LLMWhisperer v2, layout preserving output, form mode, since the packet is an image only scan and the loan application has checkboxes. 13.7 seconds for all 17 pages, 28,657 characters. LLMWhisperer puts a <<< between pages, so splitting the output into 17 page texts is one line of code.

LLMWhisperer was used to convert all pages. Showing the pages containing the pay stubs only

[PAGE 7]

09/13/2026 16:47      FROM: MARCUS FERREIRA 555-0152         TO: LAKESHORE SETUP 555-0100         P.07/17 

       EARNINGS STATEMENT . PAY DATE 08/21/2026            ADVICE ADV-20260821 

       SANDY RIVER DENTAL 
       940 Orrinbrook Avenue, Portland, OR 97233. 

        Employee                       Denise Achterberg 
        Employee Address                1980 Bellhollow Street, Portland, OR 97217 
        Pay Period                     08/03/2026 - 08/16/2026 
        Pay Date                       08/21/2026 
       Pay Frequency                   Bi-weekly 

       EARNINGS.                          HOURS       RATE          CURRENT            YTD 

        Regular                           80.00      33.19       2,655.00       45,135.00 
        GROSS : PAY                                                2,655.00       45,135:00 

        DEDUCTIONS.                                            CURRENT 
        Federal Income Tax                                       318.60 
        Oregon State Tax                                         212.40 
        Social Security                                           164.61 
        Medicare                                                  38.50 
        TOTAL DEDUCTIONS                                         734.11 

        NET PAY                                                             1,920.89 

        Direct deposit to account ending in the employee's designated account. 
        This is a system-generated earnings statement .: 

                                Sandy River Dental .. Earnings statement ADV-20260821 .. Page 1 of 1 

[PAGE 8]

09/13/2026 16:48      FROM: MARCUS FERREIRA 555-0152         TO: LAKESHORE SETUP 555-0100         P.08/17 

        EARNINGS STATEMENT . PAY DATE 08/28/2026           ADVICE ADV-20260828 

       WESTSIDE TRANSIT SERVICES 
        6100 Maristone Road, Portland, OR 97210 

       Employee                        Bartholomew Quaye 
        Employee Address               3175 Ivyfield Avenue, Portland, OR 97211 
        Pay Period                     08/10/2026 - 08/23/2026 
        Pay Date                       08/28/2026 
        Pay Frequency                  Bi-weekly 

        EARNINGS                          HOURS       RATE          CURRENT            YTD 
       Regular                            80.00      37.62       3,010.00.      51,170.00 
        GROSS .PAY.                                                3,010.00       51,170.00 

        DEDUCTIONS                                             CURRENT 
        Federal Income. Tax                                      361.20 
        Oregon State Tax                                         240.80 
        Social Security                                          186.62 
       Medicare                                                   43.65 
        TOTAL DEDUCTIONS.                                        832.27 

        NET PAY                                                             2,177.73 

        Direct deposit to account ending in the employee's designated account. 
        This is a system-generated earnings statement. 

                              Westside Transit Services . Earnings statement ADV-20260828 Page 1 of 1 

The OCR is good enough for what we are asking. There are small artefacts, for example one W-2 box reads $78.260.00 with a full stop where the comma should be, but the questions below are about what kind of page this is and whose name is on it, and none of that is affected.

The one thing we do to the text before asking Jev anything is drop the fax banner line, the one that reads “09/13/2026 16:47 FROM: MARCUS FERREIRA 555-0152 TO: LAKESHORE SETUP 555-0100 P.07/17”. It is not part of any document, and it caused a problem that we describe at the end.

On two of the 17 pages the OCR split that banner over two lines, so the code matches both halves.

The question to Jev

One call per page, three questions. A Choice for what kind of document the page belongs to, using the same vocabulary our loan intake workflow uses. A Noul for whether the page names the borrower. And a Noul for whether the page is a continuation of a document from the previous page, which is what lets the code find the boundaries between documents.

from typesafe_sdk import Choice, Noul, TypeSafeClient

BORROWER = "Bartholomew Quaye"

QUESTIONS = {
    "kind": Choice(
        instructions="What document is this page part of?",
        criteria={
            "loan_application":   "Uniform Residential Loan Application (Form 1003), "
                                  "or any one of its numbered sections",
            "photo_id":           "A driver's license, passport or state identification card",
            "pay_stub":           "An earnings statement for one pay period from an employer",
            "w2":                 "IRS Form W-2 Wage and Tax Statement for one tax year",
            "bank_statement":     "A bank's statement of an account's transactions and "
                                  "balances for a period",
            "purchase_agreement": "The contract to buy the property; a residential "
                                  "purchase agreement",
            "loan_estimate":      "The three-page Loan Estimate disclosure headed 'Loan Estimate'",
            "other":              "None of the above",
        }),
    "names_borrower": Noul(
        instructions=f"This page names {BORROWER} as the borrower, applicant, employee, "
                     f"account holder, licensee or buyer"),
    "is_continuation": Noul(
        instructions="This page continues a document that began on an earlier page, "
                     "rather than being the first page of a document"),
}

import re
FAX_BANNER = re.compile(
    r"^\s*(\d{2}/\d{2}/\d{4}\s+\d{2}:\d{2}\s+FROM:.*|(?:[A-Z]+\s+)*[\d-]+\s+P\.\d{2}/\d{2})\s*$")

def strip_fax_banner(page_text):
    return "\n".join(l for l in page_text.splitlines() if not FAX_BANNER.match(l))

for n, page_text in enumerate(pages, start=1):
    state = strip_fax_banner(page_text)
    response = TypeSafeClient().system_one(
        model="jev-1.13.0", state=state, questions=QUESTIONS)

The request for page 7, exactly as sent. Page 8’s request is the same apart from the state:

{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text of page 7, fax banner removed, 1,515 characters>",
  "questions": {
    "kind": {
      "type": "choice",
      "instructions": "What document is this page part of?",
      "criteria": {
        "loan_application": "Uniform Residential Loan Application (Form 1003), or any one of its numbered sections",
        "photo_id": "A driver's license, passport or state identification card",
        "pay_stub": "An earnings statement for one pay period from an employer",
        "w2": "IRS Form W-2 Wage and Tax Statement for one tax year",
        "bank_statement": "A bank's statement of an account's transactions and balances for a period",
        "purchase_agreement": "The contract to buy the property; a residential purchase agreement",
        "loan_estimate": "The three-page Loan Estimate disclosure headed 'Loan Estimate'",
        "other": "None of the above"
      }
    },
    "names_borrower": {
      "type": "noul",
      "instructions": "This page names Bartholomew Quaye as the borrower, applicant, employee, account holder, licensee or buyer"
    },
    "is_continuation": {
      "type": "noul",
      "instructions": "This page continues a document that began on an earlier page, rather than being the first page of a document"
    }
  }
}

What Jev returns

Page 7: 0.35 seconds, 908 input tokens. Page 8: 0.39 seconds, 911 input tokens. Both responses, verbatim:

For page 7 (Denise Achterberg’s pay stub)

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 908, "output_tokens": 124},
  "answers": {
    "kind": {
      "type": "choice",
      "choice": "pay_stub",
      "confidence": 1.0,
      "probabilities": {"pay_stub": 1.0, "loan_application": 0.0, "photo_id": 0.0, "w2": 0.0,
                        "bank_statement": 0.0, "purchase_agreement": 0.0, "loan_estimate": 0.0, "other": 0.0}
    },
    "names_borrower":  {"type": "noul", "noul": 0.01},
    "is_continuation": {"type": "noul", "noul": 0.03}
  }
}

For page 8 (Bartholomew Quaye’s pay stub)

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 911, "output_tokens": 124},
  "answers": {
    "kind": {
      "type": "choice",
      "choice": "pay_stub",
      "confidence": 1.0,
      "probabilities": {"pay_stub": 1.0, "loan_application": 0.0, "photo_id": 0.0, "w2": 0.0,
                        "bank_statement": 0.0, "purchase_agreement": 0.0, "loan_estimate": 0.0, "other": 0.0}
    },
    "names_borrower":  {"type": "noul", "noul": 0.94},
    "is_continuation": {"type": "noul", "noul": 0.03}
  }
}
QuestionPage 7Page 8
kindpay_stub 1.00pay_stub 1.00
names_borrower0.010.94
is_continuation0.030.03

Both pages are pay stubs and Jev is sure of it. Both are first pages. The only difference is the name, and Jev puts it at 0.01 against 0.94.

There is nothing uncertain about page 7. It is a perfectly good pay stub. It is just someone else’s.

We ran the same 3 questions on all 17 pages. The kind was right on every page, at 1.00 on every page. The name question was right on every page: 0.66 to 0.98 on the 13 pages that print the borrower’s name, 0.01 to 0.02 on the 4 that don’t, which are page 7 and three continuation pages that only carry the running text of a document.

The continuation question was right on every page too: 0.82 to 0.97 on the 8 pages that continue a document and 0.03 to 0.40 on the 9 that start one.

What this buys you

The extraction prompts now get one document each, of a known kind, with the schema that matches. Not 17 pages with 6 schemas mixed together. And the packet that would have put another person’s income into this borrower’s file was stopped before it cost an extraction.

This is 17 calls per packet, 700 to 1,100 tokens each depending on the page, most under a second, and they are independent so they can run in parallel. At TypeSafe’s published price the whole packet came to 15,127 tokens, about $0.0006.

What it doesn’t buy you

The first time we ran this, the continuation question was wrong on 7 of the 17 pages. The first page of the loan application, the driver license, each W-2, each bank statement and the purchase agreement all came back as a continuation, at 0.60 to 0.91. The reason was the fax banner. Every page says P.09/17 or similar at the top, and Jev reads that literally: page 9 of 17, so this page continues something. Dropping that line fixed it completely, 0 wrong out of 17. We also tried rewording the question to describe what a first page looks like and leaving the banner in. That got it down to 2 wrong. Both changes together also gave 0 wrong, with smaller margins.

The lesson is the same one as in use case 1. Jev does not know that a fax banner is not part of the document. If there is text on the page that answers your question in a way you did not intend, it will use it. Headers, footers, stamps and banners that are not part of the document should come off before you ask.

Also, this is per page. Jev never saw two pages together. Whether page 12 continues page 11 was decided from page 12 alone, from what a continuation page looks like. That worked here because bank statements and contracts look different on their first page and on their later pages. A document whose pages all look alike would need a different approach, for example asking about the pair of pages in one call.

Note: We checked whether this is a Jev-specific problem. Claude Opus 5, asked the same yes/no question with the fax banner left in, got all 17 pages right. An LLM reads a fax banner as a fax banner. Jev reads it as text, which is the price of the speed and the calibrated number, and it is why the pre-processing matters more with Jev than with an LLM. 


Use case 4: Closed-set fields: skip the LLM

The problem

A lot of what an extraction schema asks for is not free text. Is the borrower a U.S. citizen? Is this a purchase or a refinance? Will they occupy the property? Has there been a bankruptcy in the last 7 years? These are yes/no or one-of-N fields, and the options are known before you open the document, because the form itself prints them. Today these fields go to the LLM as part of one big JSON, along with the names and the numbers, and they come back with a confidence number that does not mean much.

Jev can answer them directly from the text. No prompt, no JSON to parse, a probability per field. The LLM then only has to deal with the fields that need a value written out.

The document

The loan application from the packet in use case 3. Pages 1 to 3 of the Uniform Residential Loan Application (Form 1003), from the same fax. Page 1 has the borrower’s personal information, page 2 the loan and the property, page 3 the declarations, which are a list of yes/no questions the borrower answers. It is a scanned fax like the rest of the packet.

LLMWhisperer output

Same run as use case 3: LLMWhisperer v2, layout preserving output, form mode. Pages 1 to 3 with the fax banner dropped, joined together, come to 6,843 characters. That is the state for this call. LLMWhisperer output for pages 1 to 3, before the banner lines are dropped:


09/13/2026 16:41              FROM: MARCUS FERREIRA 555-0152                       TO: LAKESHORE SETUP 555-0100                     P.01/17 

           To be completed by the Lender: 
           Lender Loan No./Universal.Loan Identifier LN-2026-040256 Agency Case No: 

           Uniform Residential Loan Application 
           Verify and complete the information on this application. If you are applying for this loan with others, each additional Borrower must provide 
           information as directed by your Lender. 

           Section 1: Borrower Information. 

           This section asks about your personal information and your income from employment and other 
           sources, such as retirement, that you want considered to qualify for this loan ... 

           1a, Personal Information 
           Name (First, Middle, Last, Suffix)        Bartholomew Quaye 
           Social Security Number                    XXX-XX-7156 
           Date of Birth (mm/dd/yyyy)                09/12/1983 
           Citizenship                               U.S. Citizen 
           Type of Credit                            Individual Credit 
           Marital Status                            Unmarried 
           Home Phone                                555-0144 
           Email                                     [email protected] 
           Current Address                           3175 Ivyfield Avenue, Portland, OR 97211 
           How long at Current Address?              3 Years 4 Months Housing: Rent 

           1b. Current Employment/Self-Employment and Income 
           Employer or Business Name                 Westside Transit Services 
           Employer Address                          6100 Maristone Road, Portland, OR 97210 
           Position or Title                         Operations Analyst 
           Time in this line of work                 5 years 4 months 
           Gross Monthly Income - Base               $6,522.00 
           Overtime / Bonus / Commission             $0.00 / $0.00 / $0.00 
           TOTAL Gross Monthly Income                $6,522.00 

           1c. IF APPLICABLE, Complete Information for Additional.Employment/Self-Employment and Income 
                                                     Does not apply 
           1d. IF APPLICABLE, Complete Information for Previous Employment/Self-Employment and Income 
                                                     Does not apply 
           1e .. Income from Other Sources 
                                                     Does not apply 

           Borrower Name: Bartholomew Quaye 
                              Uniform Residential Loan Application . Freddie Mac Form 65 . Fannie Mae Form 1003 Effective 1/2021 


<<<
555-0100          P.02/17 
09/13/2026 16:42               FROM: MARCUS FERREIRA 555-0152                      TO: LAKESHORE SETUP 

             Uniform Residential Loan Application 
             Verify and complete the information on this application. If you are applying for this loan with others, each additional Borrower must provide 
             information as directed by your Lender. 

             Section 2: Financial Information - Assets and Liabilities. 

             This section asks about things you own that are worth money and that you want considered to 
             qualify for this loan. It then asks about your liabilities (or debts) that you pay each month. 
             2a. Assets - Bank Accounts, Retirement, and Other Accounts You Have 
             Checking / Savings                        See bank statements provided with this application 
             2b. Other Assets and Credits You Have 
             Earnest Money                             Deposited under the purchase agreement 
             2c. Liabilities - Credit Cards, Other Debts, and Leases that You Owe 
             Revolving                                  One credit card account, paid monthly 

             Section 3: Financial Information - Real Estate. 

             This section asks you to list all properties you currently own and what you owe on them. 

                                                       I do not own any real estate 

             Section 4: Loan and Property Information. 

             This section asks about the loan's purpose and the property you want to purchase or refinance. 
             Loan Amount                               $373,500.00 
             Loan Purpose                              Purchase 
             Property Address                          415 Foundry Court, Portland, OR 97206 
             Number of Units                            1 
             Property Value                            $4.15,000.00 
             Occupancy                                 Primary Residence 
             4b. Other New Mortgage Loans on the PEmestyot apply 
             4c. Rental Income on the Property         Does not apply 
             4d. Gifts or Grants                       Does not apply 

             Borrower Name: Bartholomew Quaye 
                                Uniform Residential Loan Application Freddie Mac Form 65 · Fannie Mae Form. 1003 Effective 1/2021 

<<<
09/13/2026 16:43                FROM: MARCUS FERREIRA 555-0152                      TO: LAKESHORE SETUP 555-0100                      P.03/17 

             Uniform Residential Loan Application 
             Verify and complete the information on this application. If you are applying for this loan with others, each additional Borrower must provide 
             information as directed by your Lender. 

             Section 5: Declarations. 

             This section asks you specific questions about the property, your funding; and your past 
             financial history. 

             A. Will you occupy the property as your primary residence? YES 
             B. Are you borrowing any money for this real estate transaction or obtaining any money from another party ?. NO 
             C. Have you or will you be applying for a mortgage loan. on another property? NO 
             D. Are there any outstanding judgments against you? NO 
             E. Are you currently delinquent or in default on a Federal debt? . NO 
             F. Are you a party to a lawsuit in which you potentially have any personal financial liability? NO 
             G. Have you declared bankruptcy within the past 7 years? NO 

             Section 6: Acknowledgments and Agreements. 

             This section tells you about your legal obligations when you sign this application. I agree to; 
             acknowledge; and represent the following: the information I have provided is true, accurate, and 
             complete as of the date I signed this application; the Lender and others may rely on it; the 
             loan is secured by a mortgage or deed of trust on the property; my signature below applies to 
             each statement in this application. 

             Borrower Signature                         /s/ Bartholomew Quaye            Date (mm/dd/yyyy) 09/09/2026 

             Borrower Name: Bartholomew Quaye. 
                                Uniform Residential Loan Application Freddie Mac Form 65 . Fannie Mae Form 1003 Effective 1/2021 

<<<

Two things to notice. Each field is printed as a label and a value on the same line, and the value is one of a fixed set of options that the form itself lists. That is what makes these closed-set questions. And the OCR has a small artefact on page 2, $4.15,000.00 for the property value, which none of these questions touch.

The declarations end in the borrower’s YES or NO, one per line.

The question to Jev

Ten fields, one call. Five Choices, whose options are the form’s own lists, and five Nouls, one per declaration.

from typesafe_sdk import Choice, Noul, TypeSafeClient

QUESTIONS = {
    "citizenship": Choice(
        instructions="Under Personal Information, what is the borrower's Citizenship?",
        criteria={"us_citizen": "U.S. Citizen",
                  "permanent_resident_alien": "Permanent Resident Alien",
                  "non_permanent_resident_alien": "Non-Permanent Resident Alien"}),
    "type_of_credit": Choice(
        instructions="Under Personal Information, what Type of Credit is the borrower "
                     "applying for?",
        criteria={"individual": "Individual Credit", "joint": "Joint Credit"}),
    "marital_status": Choice(
        instructions="Under Personal Information, what is the borrower's Marital Status?",
        criteria={"married": "Married", "separated": "Separated", "unmarried": "Unmarried"}),
    "loan_purpose": Choice(
        instructions="Under Loan and Property Information, what is the Loan Purpose?",
        criteria={"purchase": "Purchase", "refinance": "Refinance", "other": "Other"}),
    "occupancy": Choice(
        instructions="Under Loan and Property Information, what is the Occupancy?",
        criteria={"primary_residence": "Primary Residence", "second_home": "Second Home",
                  "investment_property": "Investment Property"}),
    # Section 5, Declarations. Each line ends in the borrower's YES or NO.
    "decl_a_occupy_as_primary": Noul(
        instructions="Declaration A, 'Will you occupy the property as your primary "
                     "residence?', is answered YES"),
    "decl_b_borrowing_money": Noul(
        instructions="Declaration B, about borrowing money for this transaction or "
                     "obtaining money from another party, is answered YES"),
    "decl_c_other_mortgage": Noul(
        instructions="Declaration C, about applying for a mortgage loan on another "
                     "property, is answered YES"),
    "decl_d_judgments": Noul(
        instructions="Declaration D, 'Are there any outstanding judgments against you?', "
                     "is answered YES"),
    "decl_g_bankruptcy": Noul(
        instructions="Declaration G, 'Have you declared bankruptcy within the past "
                     "7 years?', is answered YES"),
}

state = "\n\n".join(strip_fax_banner(page) for page in pages[0:3])
response = TypeSafeClient().system_one(
    model="jev-1.13.0", state=state, questions=QUESTIONS)

The request, exactly as sent:

{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text of pages 1-3 of the loan application, fax banners removed, 6,843 characters>",
  "questions": {
    "citizenship": {
      "type": "choice",
      "instructions": "Under Personal Information, what is the borrower's Citizenship?",
      "criteria": {"us_citizen": "U.S. Citizen", "permanent_resident_alien": "Permanent Resident Alien",
                   "non_permanent_resident_alien": "Non-Permanent Resident Alien"}
    },
    "type_of_credit": {
      "type": "choice",
      "instructions": "Under Personal Information, what Type of Credit is the borrower applying for?",
      "criteria": {"individual": "Individual Credit", "joint": "Joint Credit"}
    },
    "marital_status": {
      "type": "choice",
      "instructions": "Under Personal Information, what is the borrower's Marital Status?",
      "criteria": {"married": "Married", "separated": "Separated", "unmarried": "Unmarried"}
    },
    "loan_purpose": {
      "type": "choice",
      "instructions": "Under Loan and Property Information, what is the Loan Purpose?",
      "criteria": {"purchase": "Purchase", "refinance": "Refinance", "other": "Other"}
    },
    "occupancy": {
      "type": "choice",
      "instructions": "Under Loan and Property Information, what is the Occupancy?",
      "criteria": {"primary_residence": "Primary Residence", "second_home": "Second Home",
                   "investment_property": "Investment Property"}
    },
    "decl_a_occupy_as_primary": {
      "type": "noul",
      "instructions": "Declaration A, 'Will you occupy the property as your primary residence?', is answered YES"
    },
    "decl_b_borrowing_money": {
      "type": "noul",
      "instructions": "Declaration B, about borrowing money for this transaction or obtaining money from another party, is answered YES"
    },
    "decl_c_other_mortgage": {
      "type": "noul",
      "instructions": "Declaration C, about applying for a mortgage loan on another property, is answered YES"
    },
    "decl_d_judgments": {
      "type": "noul",
      "instructions": "Declaration D, 'Are there any outstanding judgments against you?', is answered YES"
    },
    "decl_g_bankruptcy": {
      "type": "noul",
      "instructions": "Declaration G, 'Have you declared bankruptcy within the past 7 years?', is answered YES"
    }
  }
}

What Jev returns

1.37 seconds, 2,125 input tokens, 326 output tokens. The response, verbatim:

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 2125, "output_tokens": 326},
  "answers": {
    "citizenship":    {"type": "choice", "choice": "us_citizen", "confidence": 1.0,
                       "probabilities": {"us_citizen": 1.0, "permanent_resident_alien": 0.0,
                                         "non_permanent_resident_alien": 0.0}},
    "type_of_credit": {"type": "choice", "choice": "individual", "confidence": 1.0,
                       "probabilities": {"individual": 1.0, "joint": 0.0}},
    "marital_status": {"type": "choice", "choice": "unmarried", "confidence": 1.0,
                       "probabilities": {"unmarried": 1.0, "married": 0.0, "separated": 0.0}},
    "loan_purpose":   {"type": "choice", "choice": "purchase", "confidence": 1.0,
                       "probabilities": {"purchase": 1.0, "refinance": 0.0, "other": 0.0}},
    "occupancy":      {"type": "choice", "choice": "primary_residence", "confidence": 1.0,
                       "probabilities": {"primary_residence": 1.0, "second_home": 0.0,
                                         "investment_property": 0.0}},
    "decl_a_occupy_as_primary": {"type": "noul", "noul": 0.99},
    "decl_b_borrowing_money":   {"type": "noul", "noul": 0.01},
    "decl_c_other_mortgage":    {"type": "noul", "noul": 0.01},
    "decl_d_judgments":         {"type": "noul", "noul": 0.01},
    "decl_g_bankruptcy":        {"type": "noul", "noul": 0.01}
  }
}
Field Jev On the form
citizenship us_citizen, 1.00 U.S. Citizen
type_of_credit individual, 1.00 Individual Credit
marital_status unmarried, 1.00 Unmarried
loan_purpose purchase, 1.00 Purchase
occupancy primary_residence, 1.00 Primary Residence
decl_a_occupy_as_primary 0.99 YES
decl_b_borrowing_money 0.01 NO
decl_c_other_mortgage 0.01 NO
decl_d_judgments 0.01 NO
decl_g_bankruptcy 0.01 NO

10 out of 10. The five Choices at 1.00, declaration A at 0.99, the other four declarations at 0.01. Jev read each declaration’s own answer. A is YES and B to G are NO on this form, and it did not carry the YES across the block.

What this buys you

Ten fields in 1.4 seconds for about $0.0001, with no prompt to write and no JSON to parse, and a probability per field that you can threshold. The LLM now only has to deal with the fields that need a value written out: the name, the address, the employer, the loan amount, the property address. The enums and booleans never go near it. On a 1003 that is all of Section 5 and a good part of Sections 1 and 4.

What it doesn’t buy you

Position. Jev reads words. It does not see where on the page they are. On this form that does not matter, because every field is printed as a label followed by its value. On a form where the answer is a single mark in one of several columns, it matters a lot.

We tried this section first on an ACORD certificate of insurance, where the ADDL INSD and SUBR WVD columns hold a Y or nothing, and the general liability row had a Y under SUBR WVD only. Jev said ADDL INSD was marked, at 0.96. We then moved the Y to the other column in the text and asked again, and the answers did not change. It sees that there is a Y on the row. It does not see which column it is in. This seems to be a limitation on Jev’s ability to handle single characters in layout preserved text.


Use case 5: Candidate selection: several values, one field

The problem

A form prints three dates and you want one of them. It prints a policy number, an agency code, a request number and a VIN, all of which look like codes, and you want the policy number. The usual way to get these is to ask the LLM, and the LLM does two things at once: it decides which value is the right one, and it writes the value out.

Both can go wrong. It can pick the request date as the effective date, and it can transpose a digit while writing a 17 character VIN. And when it does either, it does it in the same confident JSON as everything else.

The two jobs can be separated. Code can find every value of the right shape on the page, because dates, amounts and codes have shapes a regex can match. Jev can say which one is the field. And code can copy the chosen string exactly as it was printed. Nobody writes the value, so nobody can miswrite it.

The document

A policy change request from an insurance agency, one page, asking the carrier to add a vehicle to a personal auto policy. It prints a DATE OF REQUEST at the top right, an EFFECTIVE DATE OF CHANGE in the middle, and a DATE in the signature block at the bottom.

Two of those are 09/06/2026 and one is 07/24/2026. The one that matters is the effective date, because the desk has a rule that a change cannot be backdated more than 30 days, and 07/24/2026 is 44 days before the request. Pick the wrong date and the rule passes when it should not.

The same page also prints a policy number, an agency code, a request number, a VIN, two phone numbers and three ZIP codes, and two dollar amounts, the cost of the vehicle and the deductible. Our extraction prompt for this form has rules against taking the request date as the effective date and against reading the agency code, the request number or the VIN as the policy number, because those are the mistakes we expect.

LLMWhisperer output

LLMWhisperer v2, layout preserving output, native_text mode, since the PDF is born digital. 2,539 characters. Notice how the effective date is printed: the label on one line, the value on the next.

Bluewater Insurance Services 
POLICY CHANGE REQUEST                                                                                                                       Generated by AgencyDesk · request 440187 

 PRODUCER / AGENCY                                                                       COMPANY 
 Bluewater Insurance Services                                                            Harborlight Mutual Insurance Company 
 55 Orrinbrook Street, 4th Floor                                                         PO Box 2210, Tacoma, WA 98401 · [email protected] · 555-0100 
 Brooklyn, NY 11201 
                                                                                         AGENCY CODE       BW-207                       DATE OF REQUEST      09/06/2026 
 Contact Theo Lindqvist  ·  555-0157  ·  [email protected] 

 NAMED INSURED  (name and mailing address)                                                        POLICY NUMBER 

 Priya Raghunathan                                                                                HMA-2044-0187 
 535 Gladwell Avenue, Apt 4C                                                                      LINE OF BUSINESS     Personal Auto 
 Brooklyn, NY 11215 

 CHANGE REQUESTED  (mark one) 

 [X] Add vehicle                           [ ] Add driver                            [ ] Change mailing address                [ ] Other (describe in remarks) 

 EFFECTIVE DATE OF CHANGE                             REASON FOR CHANGE 

 07/24/2026                                           Purchased vehicle 

 VEHICLE TO BE ADDED 

 YEAR     2024         MAKE      Subaru                       MODEL      Forester 

 VIN      TST4BTAFC4RH30162                                   COST NEW        $36,800 

 GARAGING ZIP       11215                                     COMP / COLL DEDUCTIBLE             $1,000                    USE   Pleasure / commute 

 REMARKS 
 Please issue the endorsement and confirm the additional premium. 

 REQUESTED BY                                               PHONE / EMAIL                                             DATE 

 Theo Lindqvist                                             555-0157  [email protected]                       09/06/2026 

This request is submitted by the producer on behalf of the insured. No change is bound until the company issues the endorsement. 
<<<

The candidates

Four regexes over the text, one per shape. The identifier regex is deliberately crude: any token of five or more characters made of capitals, digits and hyphens with at least one digit. It sweeps up phone numbers and ZIP codes along with the codes. Sorting them out is Jev’s job, not the regex’s.

import re

def found(pattern):        # every distinct match, in the order it first appears
    return list(dict.fromkeys(re.findall(pattern, text)))

CANDIDATES = {
    "dates":       found(r"\b\d{2}/\d{2}/\d{4}\b"),
    "amounts":     found(r"\$[\d,]+(?:\.\d{2})?"),
    "zips":        found(r"\b\d{5}\b"),
    "identifiers": found(r"\b(?=[A-Z0-9-]*\d)[A-Z0-9][A-Z0-9-]{4,}\b"),
}

# what that found on this page:
# dates:        09/06/2026, 07/24/2026
# amounts:      $36,800, $1,000
# zips:         98401, 11201, 11215
# identifiers:  440187, 98401, 555-0100, 11201, BW-207, 555-0157,
#               HMA-2044-0187, 11215, TST4BTAFC4RH30162

The question to Jev

Ten Choices in one call. Eight of them pick a candidate, and every one of those has “none” on the list, for the form that does not print the field. The last two are closed-set fields of the kind in use case 4, riding along because the call is already there.

from typesafe_sdk import Choice, TypeSafeClient

def pick(instructions, pool, none_text):
    criteria = {f"c{i}": value for i, value in enumerate(pool, start=1)}
    criteria["none"] = none_text
    return Choice(instructions=instructions, criteria=criteria)

questions = {
    "effective_date":  pick("Which of these is the EFFECTIVE DATE OF CHANGE requested on this form?",
                            CANDIDATES["dates"], "The effective date is not printed"),
    "date_of_request": pick("Which of these is the DATE OF REQUEST on this form?",
                            CANDIDATES["dates"], "The date of request is not printed"),
    "policy_number":   pick("Which of these is the POLICY NUMBER of the policy being changed?",
                            CANDIDATES["identifiers"], "The policy number is not printed"),
    "vin":             pick("Which of these is the VIN of the vehicle to be added?",
                            CANDIDATES["identifiers"], "No VIN is printed"),
    "agency_code":     pick("Which of these is the AGENCY CODE?",
                            CANDIDATES["identifiers"], "No agency code is printed"),
    "cost_new":        pick("Which of these is the COST NEW of the vehicle to be added?",
                            CANDIDATES["amounts"], "The cost new is not printed"),
    "deductible":      pick("Which of these is the COMP / COLL DEDUCTIBLE for the vehicle to be added?",
                            CANDIDATES["amounts"], "The deductible is not printed"),
    "garaging_zip":    pick("Which of these is the GARAGING ZIP of the vehicle to be added?",
                            CANDIDATES["zips"], "The garaging ZIP is not printed"),
    "change_type": Choice(
        instructions="Under CHANGE REQUESTED (mark one), which option is marked [X]?",
        criteria={"add_vehicle": "Add vehicle", "add_driver": "Add driver",
                  "change_mailing_address": "Change mailing address", "other": "Other"}),
    "requester_role": Choice(
        instructions="Who is submitting this request?",
        criteria={"agent": "The producer or agency, on behalf of the insured",
                  "insured": "The named insured themselves", "other": "Someone else"}),
}

response = TypeSafeClient().system_one(model="jev-1.13.0", state=text, questions=questions)

The request, exactly as sent. The three identifier questions share the same nine candidates, so the list is shown once:

{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text of the form, 2,539 characters>",
  "questions": {
    "effective_date": {
      "type": "choice",
      "instructions": "Which of these is the EFFECTIVE DATE OF CHANGE requested on this form?",
      "criteria": {"c1": "09/06/2026", "c2": "07/24/2026", "none": "The effective date is not printed"}
    },
    "date_of_request": {
      "type": "choice",
      "instructions": "Which of these is the DATE OF REQUEST on this form?",
      "criteria": {"c1": "09/06/2026", "c2": "07/24/2026", "none": "The date of request is not printed"}
    },
    "policy_number": {
      "type": "choice",
      "instructions": "Which of these is the POLICY NUMBER of the policy being changed?",
      "criteria": {"c1": "440187", "c2": "98401", "c3": "555-0100", "c4": "11201", "c5": "BW-207",
                   "c6": "555-0157", "c7": "HMA-2044-0187", "c8": "11215", "c9": "TST4BTAFC4RH30162",
                   "none": "The policy number is not printed"}
    },
    "vin":         { ... the same nine candidates, "none": "No VIN is printed" },
    "agency_code": { ... the same nine candidates, "none": "No agency code is printed" },
    "cost_new": {
      "type": "choice",
      "instructions": "Which of these is the COST NEW of the vehicle to be added?",
      "criteria": {"c1": "$36,800", "c2": "$1,000", "none": "The cost new is not printed"}
    },
    "deductible": {
      "type": "choice",
      "instructions": "Which of these is the COMP / COLL DEDUCTIBLE for the vehicle to be added?",
      "criteria": {"c1": "$36,800", "c2": "$1,000", "none": "The deductible is not printed"}
    },
    "garaging_zip": {
      "type": "choice",
      "instructions": "Which of these is the GARAGING ZIP of the vehicle to be added?",
      "criteria": {"c1": "98401", "c2": "11201", "c3": "11215", "none": "The garaging ZIP is not printed"}
    },
    "change_type": {
      "type": "choice",
      "instructions": "Under CHANGE REQUESTED (mark one), which option is marked [X]?",
      "criteria": {"add_vehicle": "Add vehicle", "add_driver": "Add driver",
                   "change_mailing_address": "Change mailing address", "other": "Other"}
    },
    "requester_role": {
      "type": "choice",
      "instructions": "Who is submitting this request?",
      "criteria": {"agent": "The producer or agency, on behalf of the insured",
                   "insured": "The named insured themselves", "other": "Someone else"}
    }
  }
}

What Jev returns

0.95 seconds, 2,029 input tokens, 582 output tokens. Every probability that is not shown is 0.00:

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 2029, "output_tokens": 582},
  "answers": {
    "effective_date":  {"type": "choice", "choice": "c2", "confidence": 1.0, "probabilities": {"c2": 1.0, ...}},
    "date_of_request": {"type": "choice", "choice": "c1", "confidence": 1.0, "probabilities": {"c1": 1.0, ...}},
    "policy_number":   {"type": "choice", "choice": "c7", "confidence": 1.0, "probabilities": {"c7": 1.0, ...}},
    "vin":             {"type": "choice", "choice": "c9", "confidence": 1.0, "probabilities": {"c9": 1.0, ...}},
    "agency_code":     {"type": "choice", "choice": "c5", "confidence": 1.0, "probabilities": {"c5": 1.0, ...}},
    "cost_new":        {"type": "choice", "choice": "c1", "confidence": 1.0, "probabilities": {"c1": 1.0, ...}},
    "deductible":      {"type": "choice", "choice": "c2", "confidence": 1.0, "probabilities": {"c2": 1.0, ...}},
    "garaging_zip":    {"type": "choice", "choice": "c3", "confidence": 1.0, "probabilities": {"c3": 1.0, ...}},
    "change_type":     {"type": "choice", "choice": "add_vehicle", "confidence": 1.0,
                        "probabilities": {"add_vehicle": 1.0, ...}},
    "requester_role":  {"type": "choice", "choice": "agent", "confidence": 1.0,
                        "probabilities": {"agent": 1.0, ...}}
  }
}
Field Candidates Jev chose Copied into the record P
effective_date 2 dates c2 07/24/2026 1.00
date_of_request 2 dates c1 09/06/2026 1.00
policy_number 9 identifiers c7 HMA-2044-0187 1.00
vin 9 identifiers c9 TST4BTAFC4RH30162 1.00
agency_code 9 identifiers c5 BW-207 1.00
cost_new 2 amounts c1 $36,800 1.00
deductible 2 amounts c2 $1,000 1.00
garaging_zip 3 ZIPs c3 11215 1.00
change_type closed set add_vehicle add_vehicle 1.00
requester_role closed set agent agent 1.00

10 out of 10, all at 1.00. The policy number came out of a list of 9 tokens that included two phone numbers and three ZIP codes. The effective date’s label is printed on the line above its value, and that did not bother Jev: 1.00 on 07/24/2026 and 0.00 on the request date.

In every case the record holds the candidate string the regex found, copied by code. Jev returned “c7”, not “HMA-2044-0187”. If the record ever holds a transposed digit, the digit was transposed on the page.

What this buys you

Two things. The choice of value and the writing of the value are now separate, and only the choice is done by a model. The writing is a string copy. And the choice comes with a probability per field. Ten fields cost about $0.0001 and took under a second. What used to be an LLM call that could pick wrong or write wrong is now a regex, a Jev call and a copy.

What it doesn’t buy you

Candidates. This works because dates, amounts and codes have a shape a regex can find. Names do not. For a supplier name or a claimant you need something to propose candidates first, which can be a list you already hold, such as your vendor master, a name finder, or an LLM. TypeSafe’s own cookbook says the same. Jev picks; it does not find.

And the arithmetic stays with you. Jev can tell you the effective date is earlier than the request date, and it did at 0.98 when we asked, but it should not be asked whether the gap is more than 30 days. Its documentation is clear that dates are text to it. Parse the strings and count the days yourself.


Use case 6: Mapping to a master list

The problem

Extracted values rarely arrive in your vocabulary. A carrier’s loss run says CLSD, another says CLOSED, a third says CWOP. Your system has five claim statuses. A supplier’s invoice line says “annual platform access” and your chart of accounts has a code for software subscriptions. Somebody has to map one onto the other, and the mapping has to be honest about the cases that do not fit. An LLM asked to map tends to force a fit. The taxonomy is closed, and “this does not map” has to be a legitimate answer.

The document

The Sentinel Bay loss run from use case 1, five claims. Each row prints a status code and a cause code in the carrier’s own vocabulary:

ClaimStatCause
SBI-556201OPENBI-AUTO
SBI-556388CLSDCOLL
SBI-556502SUBRO PENDPD-AUTO
SBI-556677CLSDGLASS
SBI-556810CLSDTHEFT

Our taxonomy has five statuses: open, closed, closed_no_payment, reopened, record_only. SUBRO PEND is not one of them. It means subrogation pending, which is a recovery process, not a claim status, and the right thing to do with it is to send it to a person, not to guess. Our cause of loss taxonomy has eleven values, and GLASS is not one of them either; the right mapping for auto glass is “other”.

LLMWhisperer output

The same text as use case 1: LLMWhisperer v2, layout preserving output, native_text mode, 2,812 characters. The Stat and Cause columns are what these questions read.

Sentinel Bay Insurance Group                                                                                                                                                                                         Page 1 
NAIC 99358  ·  Claims Services  ·  CA                                                                                                                                                             Valued as of 07/24/2026 

COMMERCIAL AUTOMOBILE — CLAIMS DETAIL 

Named Insured:   Cascade Millwork LLC 
DBA:             Cascade Cabinetry 
FEIN:            00-9100041 
Address:         17585 SW Alderfen Loop, Tualatin, OR 97062 
Policy Number:   AU-7781-0044 
Policy Period:   08/11/2025 to 08/10/2026 
Coverage:        Commercial Auto 

Claim No               DOL             Rept             Claimant / Description                                                   Cause                 Stat            Paid       O/S Reserve                Incurred 

SBI-556201             10/05/2025      10/06/2025       Vernon Pike — Rear-end collision, box truck; third-party bodily injury   BI-AUTO               OPEN        6,000.00           5,900.00              11,900.00 
SBI-556388             11/29/2025      11/30/2025       — — Backing collision at loading dock, own unit                          COLL                  CLSD        3,300.00               0.00                3,300.00 
<<<

Sentinel Bay Insurance Group                                                                                                                                                                                         Page 2 
NAIC 99358  ·  Claims Services  ·  CA                                                                                                                                                             Valued as of 07/24/2026 

COMMERCIAL AUTOMOBILE — CLAIMS DETAIL 

Claim No               DOL             Rept             Claimant / Description                                                   Cause                 Stat            Paid       O/S Reserve                Incurred 

SBI-556502             02/02/2026      02/04/2026       Ana Souza — Parking lot sideswipe, third-party property damage           PD-AUTO               SUBRO PEND 2,900.00                0.00                2,900.00 
SBI-556677             04/13/2026      04/14/2026       — — Windshield replacement, comprehensive                                GLASS                 CLSD          900.00               0.00                  900.00 
SBI-556810             06/02/2026      06/04/2026       Devon Mackey — Cargo theft from parked unit                              THEFT                 CLSD        9,400.00               0.00                8,200.00 

End of detail. Totals not provided on this report. 
<<<

The question to Jev

Ten Choices in one call, a status and a cause per claim. Each Choice lists our taxonomy, with one line of description per value, and one more option: unmapped.

from typesafe_sdk import Choice, TypeSafeClient

STATUS = {
    "open":              "The claim is open, with money still reserved",
    "closed":            "The claim is closed and payment has been made",
    "closed_no_payment": "The claim is closed and nothing was paid",
    "reopened":          "The claim was closed and has been reopened",
    "record_only":       "Recorded for information only; no payment expected",
    "unmapped":          "The printed status does not correspond to any of the above",
}
CAUSE = {
    "fire": "Fire", "water_damage": "Water damage", "weather": "Wind, hail or other weather",
    "theft_burglary": "Theft or burglary", "vandalism": "Vandalism",
    "collision": "Collision involving the insured's own vehicle",
    "liability_bodily_injury": "Liability for bodily injury to a third party",
    "liability_property_damage": "Liability for damage to a third party's property",
    "products_liability": "Products liability", "equipment_breakdown": "Equipment breakdown",
    "other": "Something else, for example glass or comprehensive",
    "unmapped": "The printed cause does not correspond to any of the above",
}

questions = {}
for cn in ["SBI-556201", "SBI-556388", "SBI-556502", "SBI-556677", "SBI-556810"]:
    k = cn.replace("-", "_")   # hyphens out of the question names
    questions[f"{k}__status"] = Choice(
        instructions=f"The Stat (status) printed for claim {cn}, mapped to the standard claim status",
        criteria=STATUS)
    questions[f"{k}__cause"] = Choice(
        instructions=f"The Cause printed for claim {cn}, mapped to the standard cause of loss",
        criteria=CAUSE)

response = TypeSafeClient().system_one(
    model="jev-1.13.0", state=whisperer_text, questions=questions)
The request, exactly as sent, showing the first claim's two questions. The other eight are the same with the claim number changed:
{
  "model": "jev-1.13.0",
  "state": "<the LLMWhisperer text of the loss run, 2,812 characters>",
  "questions": {
    "SBI_556201__status": {
      "type": "choice",
      "instructions": "The Stat (status) printed for claim SBI-556201, mapped to the standard claim status",
      "criteria": {
        "open": "The claim is open, with money still reserved",
        "closed": "The claim is closed and payment has been made",
        "closed_no_payment": "The claim is closed and nothing was paid",
        "reopened": "The claim was closed and has been reopened",
        "record_only": "Recorded for information only; no payment expected",
        "unmapped": "The printed status does not correspond to any of the above"
      }
    },
    "SBI_556201__cause": {
      "type": "choice",
      "instructions": "The Cause printed for claim SBI-556201, mapped to the standard cause of loss",
      "criteria": {
        "fire": "Fire",
        "water_damage": "Water damage",
        "weather": "Wind, hail or other weather",
        "theft_burglary": "Theft or burglary",
        "vandalism": "Vandalism",
        "collision": "Collision involving the insured's own vehicle",
        "liability_bodily_injury": "Liability for bodily injury to a third party",
        "liability_property_damage": "Liability for damage to a third party's property",
        "products_liability": "Products liability",
        "equipment_breakdown": "Equipment breakdown",
        "other": "Something else, for example glass or comprehensive",
        "unmapped": "The printed cause does not correspond to any of the above"
      }
    },
    "SBI_556388__status": { ... },
    "SBI_556388__cause":  { ... },
    "SBI_556502__status": { ... },
    "SBI_556502__cause":  { ... },
    "SBI_556677__status": { ... },
    "SBI_556677__cause":  { ... },
    "SBI_556810__status": { ... },
    "SBI_556810__cause":  { ... }
  }
}

What Jev returns

0.91 seconds, 3,126 input tokens, 1,066 output tokens. The response, with probabilities of 0.00 left out; everything shown is as returned:

{
  "model": "jev-1.13.0",
  "usage": {"input_tokens": 3126, "output_tokens": 1066},
  "answers": {
    "SBI_556201__status": {"type": "choice", "choice": "open", "confidence": 1.0,
                           "probabilities": {"open": 1.0, ...}},
    "SBI_556201__cause":  {"type": "choice", "choice": "liability_bodily_injury", "confidence": 0.87,
                           "probabilities": {"liability_bodily_injury": 0.88, "collision": 0.07, "unmapped": 0.05, ...}},
    "SBI_556388__status": {"type": "choice", "choice": "closed", "confidence": 1.0,
                           "probabilities": {"closed": 1.0, ...}},
    "SBI_556388__cause":  {"type": "choice", "choice": "collision", "confidence": 1.0,
                           "probabilities": {"collision": 1.0, ...}},
    "SBI_556502__status": {"type": "choice", "choice": "unmapped", "confidence": 0.82,
                           "probabilities": {"unmapped": 0.85, "open": 0.14, "closed": 0.01, ...}},
    "SBI_556502__cause":  {"type": "choice", "choice": "liability_property_damage", "confidence": 0.94,
                           "probabilities": {"liability_property_damage": 0.95, "unmapped": 0.03,
                                             "other": 0.01, "collision": 0.01, ...}},
    "SBI_556677__status": {"type": "choice", "choice": "closed", "confidence": 1.0,
                           "probabilities": {"closed": 1.0, ...}},
    "SBI_556677__cause":  {"type": "choice", "choice": "other", "confidence": 0.98,
                           "probabilities": {"other": 0.99, "unmapped": 0.01, ...}},
    "SBI_556810__status": {"type": "choice", "choice": "closed", "confidence": 1.0,
                           "probabilities": {"closed": 1.0, ...}},
    "SBI_556810__cause":  {"type": "choice", "choice": "theft_burglary", "confidence": 1.0,
                           "probabilities": {"theft_burglary": 1.0, ...}}
  }
}
ClaimPrintedJevconfCode does
SBI-556201OPENopen 1.001.00write
SBI-556201BI-AUTOliability_bodily_injury 0.88, collision 0.070.87write
SBI-556388CLSDclosed 1.001.00write
SBI-556388COLLcollision 1.001.00write
SBI-556502SUBRO PENDunmapped 0.85, open 0.140.82review, raw: SUBRO PEND
SBI-556502PD-AUTOliability_property_damage 0.950.94write
SBI-556677CLSDclosed 1.001.00write
SBI-556677GLASSother 0.990.98write
SBI-556810CLSDclosed 1.001.00write
SBI-556810THEFTtheft_burglary 1.001.00write

All ten as our own lookup table says. The two worth looking at are the two that are not 1.00. SUBRO PEND went to unmapped at 0.85 with 0.14 left on open, which is a fair reading: a claim in subrogation is not closed, but subrogation pending is not a status in our list. Nobody forced it into open. And BI-AUTO on a rear-end collision with a third party injured split 0.88 for liability, bodily injury against 0.07 for collision. The description mentions a collision; the code says bodily injury. Jev leaned the way our table does and kept a little on the other reading.

What this buys you

A mapping that says no. The unmapped option is doing real work here: the one code that should not be forced into the taxonomy was not. And a probability per mapping, so the threshold is yours to set. Ten mappings cost about $0.00013 and took under a second.

The same shape works for anything that has to land in a list you own. Line items into a chart of accounts. Supplier names into a vendor master. A Choice holds up to 255 options, which covers most taxonomies outright. For a longer list, a chart of accounts with a thousand codes or a vendor master with ten thousand names, shortlist first with a fuzzy match and let Jev pick from the shortlist, with none on the list.

What it doesn’t buy you

The mapping itself is still your decision. Jev does not know whether your business treats a claim in subrogation as open or closed. It told you the wording does not fit the list, and put 0.14 on open, which is as much as it should say. Someone still has to decide, and the decision belongs in the lookup table, not in the model.

And the numbers move a little between calls. We ran these ten questions twice, once as a feasibility check and once for this section. BI-AUTO came out 0.74 liability against 0.18 collision the first time and 0.88 against 0.07 the second. Nothing else moved by more than 0.02. A threshold of 0.7 wrote it both times; a threshold of 0.8 would have sent the first run to review. Do not put a threshold on a knife edge, and expect the borderline cases to be the ones that move.