A few months ago, I argued that the future of AI in critical systems lies less in open‑ended chatbots and more in deterministic, rule‑constrained models. At the time, that was mostly a conceptual claim. With the recent launch of Jev, from an AI lab called TypeSafe AI, that thesis now has a concrete commercial example. Jev is presented not as another chatbot, but as a ‘smart if-then-else ‘ construct that returns typed, schema‑bound (having predefined structure and set of allowed values) decisions instead of free text. Moving from text generation to decision making has direct implications for how we design, procure, and govern AI in various domains. It also raises a simple but uncomfortable question: if many real‑world tasks are really about classification and choice, why do we keep reaching for models whose main skill is writing?
The newsletter Stacked recently ran a simple audit question: ‘I need to classify and route customer support tickets automatically. What should I use?’ Five different AI assistants were asked this question. Almost all of them treated it as a prompt to recommend LLM‑based solutions or large SaaS platforms. Only a couple mentioned traditional machine learning methods like fine‑tuned classifiers or lightweight natural language processing (NLP) pipelines, and even those answers quickly drifted back to agents and function calling. The task itself is straightforward. A customer writes that they were billed twice. The system needs to decide: is this a payment issue, a delivery issue, or a return? Is a refund wanted? Based on that, the ticket is routed to the right team. No essay or conversation is needed. You need a fast, reliable label and a small set of structured fields. Now imagine a similar case in a power plant, a water utility, or a national transport network. A sensor stream shows an anomaly. The system must decide: normal, warning, or shut down now. A citizen reports a fault in public infrastructure. The system must decide: which department, which priority level, which response time. These are decision tasks, not chat tasks. TypeSafe’s Jev is built around this distinction. Instead of generating text token by token, it takes structured input and a fixed schema of allowed answers, then returns labels and numbers with associated confidence scores. There are no output tokens, because there is no generated text to charge for or parse. TypeSafe describes it as a ‘function call’ rather than a chatbot: unstructured state in, typed probabilistic decisions out. In practical terms, Jev is given a set of questions and a list of permitted answers. For the billing example, the schema might allow only three categories (payment, delivery, return) and a yes/no field for “refund wanted”. Jev cannot invent a fourth category. It can still pick the wrong one, but it cannot step outside the defined answer space. The confidence score tells your system how much to trust that decision and whether to escalate to a human. This is why TypeSafe calls Jev a ‘smart if‑then‑else‘ construct. It is a component you can embed wherever a handwritten rule is too brittle, but a full LLM call is too slow, too expensive, or too open‑ended. The idea is not entirely new. Researchers have explored similar designs before, including a solo project called SalesRLAgent that also skipped token generation and returned a single probability output in milliseconds. Jev is notable because it packages this approach as a general‑purpose decision engine and has attracted significant attention. It is natural to ask: if the outputs are limited to a fixed list of options, what makes this “AI” and not just a glorified decision tree? The difference lies in how the system creates the mapping from input to output. In traditional rule‑based systems, engineers manually write conditions: ‘if the message contains ‘refund’ and ‘not received’, then category = delivery’. Such systems can handle only cases someone explicitly anticipated. Any unusual phrasing or new situation breaks the flow. With Jev‑style models, humans define the allowed answers, but not the rules that lead to them. The model is trained on data (for example, past support tickets or sensor readings) and learns complex patterns that correlate with each label. It can generalise to new phrasings and edge cases it has never seen, and it returns confidence scores that reflect how certain it is. The schema constrains the output space, not the reasoning process that leads to a choice. In other words, the interface may look similar to old support bots (pick from a few options), but the engine under the hood is a learned, probabilistic model rather than a hand‑coded decision tree. This is why it is treated as AI, even though it cannot invent new categories beyond what we defined. In industrial and infrastructure settings, many AI use‑cases are fundamentally about classification and choice: For these tasks, deterministic, schema‑bound outputs are preferable to free‑form text. Natural language should not be parsed in a safety loop. An AI tool should not be expected to ‘explain’ its reasoning in a paragraph when what is needed is a label and a confidence score. The interface should resemble the rest of the control logic: inputs, outputs, types, thresholds. Jev‑style models fit this approach. They return results in 70–500 ms, compared with 3–329 seconds for frontier LLMs in the same newsletter’s benchmarks. They charge only for input tokens because they generate no output tokens. That speed and simplicity make them suitable for high‑volume, low‑latency workflows where an LLM would be excessive. This is the structural turn in practice. Instead of asking an AI to ‘talk about’ a situation, we ask it to choose from a fixed set of options and to attach a calibrated confidence to that choice (enforcing a deterministic output schema over a probabilistic model). The system remains constrained, auditable, and integrable with existing engineering patterns.
The efficiency gains are not just a technical nicety. They have direct policy and development implications. Because decision‑only models do not generate text, they use far fewer tokens per call. TypeSafe reports response times of 70–500 ms and no output token charges, unlike LLMs that both read and generate many tokens. For high‑volume tasks, such as processing millions of support tickets, sensor events, or maintenance requests, the cumulative savings in compute and cost can be substantial. Lower token usage means lower cloud bills. That makes it easier for smaller institutions, municipalities, and agencies to run AI pilots without locking themselves into expensive, long‑term contracts with a handful of large vendors. It also means less compute per decision, which in most data centres translates into less energy use. As governments and organisations adopt AI and sustainability goals, choosing decision-only models for high-volume, low-complexity tasks becomes a concrete way to reduce AI’s carbon footprint without sacrificing functionality. In other words, choosing the right type of AI for the task is both a technical optimisation and a governance and development choice. Procurement guidelines that default to ‘use an LLM’ for every AI task will tend to push public bodies toward more expensive, more energy‑intensive solutions, even when simpler, faster, and cheaper options exist. Many AI for development (AI4D) use‑cases are classification and routing tasks: For these tasks, frontier LLMs are often the wrong tool. They are expensive, energy‑intensive, and designed for open‑ended generation. Decision‑only models are a better fit. They are faster, cheaper, and easier to integrate with existing digital systems. They can run on modest hardware and with smaller cloud budgets. Encouraging decision‑engine systems in development projects could reduce dependence on a few large model providers. It would make AI pilots more sustainable over time, because ongoing token costs are lower. It would also allow local teams to build and maintain their own classifiers for domain‑specific tasks, using open tools and datasets. This aligns with broader goals of digital sovereignty and capacity building. Instead of importing black-box chatbots for every problem, countries can develop a portfolio of lightweight, purpose-built decision models for high-volume operational tasks. Current AI laws and guidelines mostly talk about ‘foundation models’ or ‘general‑purpose AI ‘. They focus on models that generate text, images, or code. They still lack a clear category for embedded decision engines that return only labels and numbers, creating a governance blind spot. Consider questions like: Who is accountable when a calibrated decision model misclassifies a critical event in a power grid or a hospital? Or, how do auditors inspect systems that do not produce human‑readable reasoning, only labels and confidence scores? Today, many frameworks treat a Jev‑style system as just another AI component, without recognising that its risk profile differs from that of an open‑ended chatbot. The failure modes are not hallucinated paragraphs or offensive text. They are systematic misclassifications, miscalibrated confidence scores, and opaque training data that biases certain decisions. It is still early to prescribe detailed rules for decision‑only AI. For now, it is enough to recognise that this category exists and will need its own treatment in standards, audits, and tender documents. The next section sketches some concrete directions for how policymakers and procurers can start responding.
Policymakers and public procurers can start treating decision‑only AI as a distinct category with some concrete steps: Specify whether the task requires free‑form text generation or structured classification and choice. Ask whether the system produces free text or fixed‑schema outputs, and what the expected token usage per decision is. Favour solutions that minimise unnecessary generation and compute, and request estimates of energy implications. Support projects that use fine-tuned traditional models, spaCy pipelines, or decision-only engines for classification and routing, reserving frontier LLMs for tasks that truly need open-ended reasoning. Fund public datasets for fault detection, alarm triage, and service routing, so that smaller players can compete and audit. These steps do not require new laws. They can be implemented through guidance notes, template tender documents, and pilot programmes. The next wave of AI in critical systems will not look like a chatbot. It will look like invisible logic that decides. It will sit inside control rooms, ticketing systems, and maintenance workflows, returning labels and numbers instead of paragraphs. In systems like these, we cannot rely on open‑ended probabilistic text. We need constrained, auditable decisions with known failure modes and tied to clear accountability. Jev is one early sign of this move. It shows commercial demand for AI that does not talk, but decides. It also shows that our governance frameworks are not yet designed for this reality. If the future of AI in infrastructure and public services is indeed deterministic, as I argued before, then decision‑only models are among the first signs that this future is already being built. The task for policymakers, procurers, and development agencies is to recognise this trend early and shape it so it serves public interest, not just vendor roadmaps. That means writing rules, tenders, and pilot programmes that treat decision‑only AI as a distinct category, with its own risks, benefits, and requirements. Author: Slobodan Kovrlija
The ticket‑routing question
What Jev actually is
Significance for industry and infrastructure

Efficiency as a policy and development concern
Implications for development projects
The governance blind spot
Practical pointers for policymakers
A deterministic future, in practice