<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/">
  <channel>
    <title>The Uprelic Blog</title>
    <link>https://platform.uprelic.com/blog</link>
    <atom:link href="https://platform.uprelic.com/blog/feed.xml" rel="self" type="application/rss+xml" />
    <description>Notes on building Strom, the decision API for business automation.</description>
    <language>en</language>
    <lastBuildDate>Sat, 03 Oct 2026 08:00:00 GMT</lastBuildDate>
    <item>
      <title>Why we built Strom, an EU alternative decision model</title>
      <link>https://platform.uprelic.com/blog/strom-eu-alternative-decision-model-to-jev</link>
      <guid isPermaLink="true">https://platform.uprelic.com/blog/strom-eu-alternative-decision-model-to-jev</guid>
      <pubDate>Sat, 03 Oct 2026 08:00:00 GMT</pubDate>
      <dc:creator>Marco Herzog</dc:creator>
      <category>Company</category>
      <description>Small decision models lose accuracy. Large ones need datacenter GPUs and heavy post-training. None of the open models matched Jev overall, and Jev can't see images. Our data had to stay in the EU. So we trained our own, and now it's an API.</description>
      <content:encoded><![CDATA[<p>Most of you have probably played with open-source decision models by now, or built your own. Since Jev launched on September 15, the idea of a model that only makes decisions (yes or no, a score, one option out of many) has spread faster than anything we&#39;ve seen. By one count, Hugging Face had more than 1,300 decision-model repositories eighteen days later.</p>
<p>We were building on these models too. We kept running into the same problems.</p>
<h2 id="what-we-ran-into">What we ran into</h2>
<h3 id="small-models-lose-accuracy">Small models lose accuracy</h3>
<p>Decision models are cheap to clone, and the small ones are tempting because they run on a laptop. But accuracy falls off quickly below a few billion parameters. The kev family published the clearest curve so far, on its own test suite.</p>
<figure><img src="https://platform.uprelic.com/blog/media/strom-eu-alternative-decision-model-to-jev/size-curve.svg" alt="Bar chart of test accuracy by model size for the kev family, with 69.7% at 0.8B, 83.8% at 4B, 85.2% at 9B and 88.9% at 27B parameters" /><figcaption>Accuracy rises fastest up to 4B. From 4B to 27B it&#39;s another 5.1 points, which cuts the errors by about a third. Data from the kev family&#39;s new-source test split.</figcaption></figure>
<p>The smallest model is wrong almost one time in three, and its probabilities are much less reliable, with a Brier score of 0.416 against 0.156 at 27B (lower is better). Both matter when you want to automate a decision. A wrong answer with a confident probability is worse than no answer.</p>
<p>The bigger models keep paying off where it counts for automation. From 4B to 27B, the error rate drops from 16% to 11%, the Brier score from 0.242 to 0.156, and only the 27B model holds its accuracy on inputs up to 64K tokens, against 8K for the smaller ones. Part of that gap is the method, since the 27B model was fully fine-tuned and the smaller ones with LoRA adapters.</p>
<h3 id="large-models-need-datacenter-gpus-and-heavy-post-training">Large models need datacenter GPUs and heavy post-training</h3>
<p>The 27B class and bigger is accurate, but it runs on datacenter GPUs. The base weights alone don&#39;t get you there. One small model, Laya, scores 36% on a decision benchmark before fine-tuning and 77% after. In decision models, the fine-tune is the model. Every team that wants the accuracy has to pay for the training, and then for the GPUs.</p>
<p>Size also sets how far post-training can take a model. In a 2026 study of supervised fine-tuning on models up to 235B parameters, bigger models ended up clearly better after the same training, with diminishing returns, and more so with full fine-tuning than with LoRA adapters. More training examples helped every size by about the same fraction (<a href="https://arxiv.org/abs/2609.01244">O&#39;Neill et al., 2026</a>; see also <a href="https://arxiv.org/abs/2402.17193">Zhang et al., 2024</a>). Larger models also need less data to reach their best. On synthetic training data, an 8B model peaked at 1 trillion tokens, a 3B model only at 4 trillion (<a href="https://arxiv.org/abs/2503.19551">Qin et al., 2025</a>). For a decision model whose training set keeps growing, that&#39;s the point because a bigger base keeps its lead with every round of new data.</p>
<p>Most open-source projects don&#39;t use extensive training data, usually don&#39;t clean it enough, and don&#39;t enrich it with synthetic data. To be fair, for open source this is too expensive. To get a model that performs well, the whole data pipeline has to be an investment, with continuous training. For us, that means we serve one model while training several others, and we&#39;re already preparing the data for the second generation.</p>
<h3 id="jev-can-t-see-images">Jev can&#39;t see images</h3>
<p>First of all, a big shout-out to TypeSafe AI, the creators of Jev. They opened up a completely new class of models, and excitement for new use cases. They truly deserve respect for the idea and for the many new approaches that will come with it.</p>
<p>One big missing piece, and a crucial one for us, is that Jev doesn&#39;t take images. For many of our cases that&#39;s a must. Sometimes an image is the fallback when the text isn&#39;t enough, like the photo attached to an insurance damage report or the screenshot in a support ticket. Sometimes there&#39;s no text at all, as with a product photo that has to meet listing rules or a design question. You get the point.</p>
<p>Images are just as important for our upcoming extension, a browser-use agent. A web agent also decides based on what it sees, like which button to click, whether a form was filled in correctly, whether a page shows an error and needs a visual fallback, and what to do when there&#39;s no DOM at all, as on a canvas. Sometimes it simply looks to double-check. The page&#39;s text often leaves out what matters, so the screenshot is part of the state. A decision model that can&#39;t read an image can&#39;t drive the agent reliably.</p>
<figure class="post-video"><video src="https://platform.uprelic.com/blog/media/strom-eu-alternative-decision-model-to-jev/strom-vs-jev.mp4" poster="https://platform.uprelic.com/blog/media/strom-eu-alternative-decision-model-to-jev/strom-vs-jev.jpg" controls muted playsinline preload="metadata" aria-label="Strom and Jev side by side on Google Flights, driving the same browser agent toward the same round trip from Zürich to New York"></video><figcaption>The same agent code and goal, started together; only the model that decides is different. Both get the same instruction for date pickers, which is to click the field, then the date, then confirm. After the departure date, Google keeps the calendar open with the return field already active. Jev follows the instruction literally and clicks the field again and again, while Strom picks October 27 straight away, from the same page text and without a screenshot. Strom finds the flights in 7.7 seconds. Across all our recorded runs of this task, Strom passed 20 of 20 and Jev 1 of about 12. It shows one situation the agent&#39;s instructions don&#39;t cover, not a general ranking.</figcaption></figure>
<p>In upcoming posts, we&#39;ll take these image use cases one by one, from damage photos in insurance claims and screenshots in support tickets to product photos against listing rules and the browser agent.</p>
<h3 id="and-our-data-had-to-stay-in-the-eu">And our data had to stay in the EU</h3>
<p>Many of the decisions we automate run on personal data, like customer emails, invoices and photos. Under the GDPR, every provider that processes that data for you needs a data processing agreement, and every transfer outside the EU needs its own safeguards. So the shortest path to compliance was our own model on GPUs in Paris, run for us by Scaleway, a French cloud provider, with no third-party AI provider in between. When a customer&#39;s data protection officer asks where their data goes, we wanted the answer to be simple. It&#39;s processed in France.</p>
<h2 id="so-we-trained-our-own">So we trained our own</h2>
<p>Strom is a large vision-language model, fine-tuned on business decisions in currently 14 domains, from insurance and finance to mechanical engineering and office work, with tasks from invoices and contracts to machine logs and web agents. It takes free-form text, JSON or images, and answers named questions with one of the options you define, plus a probability for each option. It speaks the same request format as Jev, with <code>choice</code>, <code>score</code> and <code>noul</code> questions. For its size and precision it needs datacenter-class GPUs.</p>
<h2 id="why-it-s-an-api-now">Why it&#39;s an API now</h2>
<p>We mostly built Strom for ourselves, for <a href="https://uprelic.com">uprelic.com</a>, our main service, which is a co-work chat. In upcoming posts we&#39;ll share how this model class can improve chat applications too, because some approaches, like model routing, sound nice in theory but aren&#39;t feasible yet. At the same time, it was a no-brainer that other teams have the same needs, like decisions they want to automate, data that should stay in the EU, inputs that are often images and a dirt-cheap workflow pipeline, but no capacity to build and continuously improve a model themselves.</p>
<p>So we&#39;re opening it up. Strom costs <strong>$0.042 per million input tokens, and output is free</strong>, the same per token as Jev.</p>
<h2 id="how-strom-performs">How Strom performs</h2>
<p>One caution before the numbers. A model&#39;s own test suite says little about how it compares with others. kev-27B scores 88.9% on kev&#39;s own suite, ahead of Jev&#39;s 85.7% there, yet no open model has matched Jev overall. Teams build their tests around what their model was trained on, often without meaning to. That applies to us too. So we stick to public benchmarks, point out where our runs aren&#39;t fully clean, and keep improving how we measure, with more held-out and independent benchmarks over time.</p>
<p>We measured Strom 1.0.7 in production on October 1, 2026, on public benchmarks whose results for other models are published. Here is what we found, including where Strom is behind.</p>
<h3 id="on-par-with-jev-on-jevbench">On par with Jev on JevBench</h3>
<p>On the 231 public JevBench tasks, Strom gets 201 right. Jev 1.13 gets 200, according to the published results.</p>
<table>
<thead>
<tr>
<th>System</th>
<th align="right">Correct (of 231)</th>
<th align="right">Brier</th>
<th align="right">ECE</th>
</tr>
</thead>
<tbody><tr>
<td>GPT-6 Astra (reasoning, low)</td>
<td align="right">231 (100%)</td>
<td align="right">0.009</td>
<td align="right">0.015</td>
</tr>
<tr>
<td>GPT-5.6 Luna (no reasoning)</td>
<td align="right">206 (89.2%)</td>
<td align="right">0.207</td>
<td align="right">0.093</td>
</tr>
<tr>
<td><strong>Strom 1.0.7</strong></td>
<td align="right"><strong>201 (87.0%)</strong></td>
<td align="right"><strong>0.167</strong></td>
<td align="right"><strong>0.033</strong></td>
</tr>
<tr>
<td>Jev 1.13.0</td>
<td align="right">200 (86.6%)</td>
<td align="right">0.181</td>
<td align="right">0.032</td>
</tr>
<tr>
<td>Open-Jev 27B v1.1</td>
<td align="right">197 (85.3%)</td>
<td align="right">0.242</td>
<td align="right">–</td>
</tr>
<tr>
<td>Open-Jev 9B</td>
<td align="right">179 (77.5%)</td>
<td align="right">0.322</td>
<td align="right">0.086</td>
</tr>
<tr>
<td>Open-Jev 2B</td>
<td align="right">150 (64.9%)</td>
<td align="right">0.475</td>
<td align="right">0.127</td>
</tr>
</tbody></table>
<p>The reasoning model gets everything right, at a median of 2.2 seconds per answer. Strom answered in a median of 136 ms, measured from a client in Germany, network included. Latencies in the published results come from other clients, so only the accuracy and calibration columns compare directly.</p>
<p>Strom&#39;s probabilities are well calibrated. An expected calibration error (ECE) of 0.033 means its stated confidence is, on average, about three points away from how often it&#39;s actually right. That&#39;s what makes thresholds work. If 0.9 means right nine times in ten, your code can act on it.</p>
<p>Strom gets every task right in extraction, intent, routing, policy, tool selection and ordinal scoring. It&#39;s weakest on dates and numbers (5 of 15). It also struggles with long policies (12 of 19) and questions that chain several facts (13 of 18). We&#39;re addressing these in our next training cycle.</p>
<h3 id="where-jev-is-ahead">Where Jev is ahead</h3>
<p>On the control suites Open-Jev published, Jev is ahead.</p>
<table>
<thead>
<tr>
<th>Suite</th>
<th align="right">Strom 1.0.7</th>
<th align="right">Jev 1.13.0</th>
</tr>
</thead>
<tbody><tr>
<td>FizzBuzz (300)</td>
<td align="right">264</td>
<td align="right">299</td>
</tr>
<tr>
<td>Mailroom (921)</td>
<td align="right">898</td>
<td align="right">908</td>
</tr>
<tr>
<td>JF100 (300)</td>
<td align="right">209</td>
<td align="right">232</td>
</tr>
</tbody></table>
<p>The errors have a pattern. In FizzBuzz, Strom never misses &quot;divisible by 5&quot;, but says yes to &quot;divisible by 3&quot; for numbers like 13 and 29. Every Mailroom error is a false yes. In JF100, relations, dates and arithmetic account for most misses, while policy rules and discourse are perfect. Here too, we&#39;re fixing these in the training data, one family at a time.</p>
<h3 id="an-early-signal-on-images">An early signal on images</h3>
<p>Image JevBench publishes questions and answers, not images, so we rebuilt part of the image set ourselves. On its 128 public preview items, Strom gets 94 right (73.4%). The best system in the benchmark&#39;s own preview round got 59.4%. Treat that as a sanity check, not a score, because 68 of the images are ours, and some of our versions are probably easier than the originals. A real comparison needs a run on the benchmark&#39;s sealed images.</p>
<h3 id="speed">Speed</h3>
<p>Time in the model depends mostly on how much you send. The chart shows medians of 10 requests per point, with fresh content each time, measured in production.</p>
<figure><img src="https://platform.uprelic.com/blog/media/strom-eu-alternative-decision-model-to-jev/speed.svg" alt="Line chart of median time in the model against input tokens, where text rises from 91 ms at 550 tokens to 228 ms at 8,000 and 673 ms at 24,000, and one image rises from 58 ms to 168 ms between 256 and 1024 pixels wide" /><figcaption>A yes-or-no question on 550 tokens of text takes 91 ms, on 8,000 tokens 228 ms; one image takes 58 ms at 256 pixels wide and 168 ms at 1024. Production, October 1, 2026.</figcaption></figure>
<p>Add your network round trip, which is about 100 ms from Germany. The <a href="https://platform.uprelic.com/docs/latency">latency guide</a> has the full measurements, including how more questions and options add up. We plan to run in more locations to bring latency down.</p>
<h2 id="what-we-re-betting-on">What we&#39;re betting on</h2>
<p>Eighteen days in, decision models are a category. The format has been rebuilt by an open lab, a runtime, a CDN and hundreds of independent authors. Weights alone won&#39;t be the moat for anyone. What&#39;s left is what you can&#39;t get from a weekend fine-tune, namely calibrated probabilities you can set thresholds on and a continuous data pipeline on infrastructure that processes your data in the EU.</p>
<p>That&#39;s what Strom is for. The <a href="https://platform.uprelic.com/docs">docs</a> show how to send your first request.</p>
<hr>
<h3 id="data-and-method">Data and method</h3>
<p>We measured Strom 1.0.7 in production on October 1, 2026. For JevBench we used the 231 public tasks with the Open-Jev harness&#39;s request format and scoring, and took the results of other systems from the Open-Jev benchmarks page. We rebuilt the control suites (FizzBuzz, Mailroom, JF100) the way Open-Jev built them and checked all 1,521 gold labels against Open-Jev&#39;s published rows. Two items in our training data turned out to resemble JevBench tasks, one long policy and one date question, so those two slices aren&#39;t fully clean for Strom. The kev numbers come from the kev family&#39;s published evaluation (new-source test split), the Laya numbers from Convai Innovations&#39; published model card, and the ecosystem count from Hugging Face repositories with decision-model tags created after September 15, 2026.</p>
]]></content:encoded>
    </item>
  </channel>
</rss>
