← Back to Blog

Testing Jev on MentalBench

• By John Britton

TypeSafe recently released Jev, which it calls its first System One model. System One comes from Daniel Kahneman's idea of fast System 1 versus more deliberate System 2 thinking. So it's a nonthinking model that works quickly. Jev is a general classifier that doesn't need fine-tuning. It doesn't give back text like a chatbot; it takes information and answers questions about it. You can ask it true-or-false questions, ask it to choose among categories, or ask for a score, all just using natural language. TypeSafe says it's a model built for fast decisions, like a decision layer inside software. Another interesting thing about Jev is that it gives a probability score for each answer. A person or the software could decide what to do when something is flagged at a certain probability level.

What does all this have to do with psychology? A lot of psychology is built on language and discussion between people, but also concepts that oftentimes are not quantitative in nature. Being able to take that type of unstructured information and get structured research insights from it at scale is something that Jev could potentially do really well. I wanted to show a proof-of-concept project on a public dataset, with a real example of what it costs to do this kind of research. MentalBench is the example I found with ChatGPT's help. After the project, I'll explore some other possibilities, including real-time use in a conversation or transcript. Each of those applications would need its own testing.

Why MentalBench?

MentalBench is a publicly available benchmark with 24,750 synthetic psychiatric case descriptions covering 23 diagnoses. The dataset is available on Hugging Face, where its card identifies it as CC BY 4.0. It was built from a knowledge graph of DSM-5 criteria and differential diagnosis rules. The authors worked with clinical experts on the knowledge graph, and a psychiatrist and a clinical psychologist independently reviewed 220 of the synthetic cases for diagnostic validity, realism, and naturalness. Even though these cases are synthetic, that expert review gives us some evidence about their quality and makes this a useful benchmark for a test like this one. It doesn't validate every case or establish accuracy on real patients.

The dataset has four types of cases. Type 1 includes case write-ups as if written by a medical professional, with one correct diagnosis. Type 2 is closer to how a patient might tell the story, also with one correct diagnosis. In Type 3, two different differential diagnoses might be plausible because information is missing, so two diagnoses are accepted. Type 4 gives the information needed to decide which one fits, and has one correct diagnosis. Each case supplies four possible diagnoses. The original paper tested LLMs on those four options and scored an answer as correct only when the complete answer set matched.

I decided to test Jev in two ways over this dataset. I wanted to compare it against the original way. The MentalBench paper tested the dataset with only four possible diagnoses, and I wanted to also test it against all 23 diagnoses, which is a higher standard. I thought Jev would do a good job of keeping track of all that structured information, and the 23-diagnosis version would tell us more about that capability. I was hoping for strong accuracy with the 23, but the project itself, and being able to conduct it over a large dataset, was an overarching goal. My hypothesis was that Jev might perform as well as or better than the best LLM reported in the paper, at a much lower cost.

There are privacy and ethical considerations here. This project used publicly available synthetic cases. Working with real patient information would require appropriate permissions and privacy arrangements. In a HIPAA-covered setting where a cloud service handles protected health information, that includes a business associate agreement and other safeguards. I couldn't find a public indication that TypeSafe currently offers a BAA for Jev, though someone considering that use would need to confirm directly with TypeSafe. TypeSafe's privacy policy says it does not train or fine-tune on customer inputs, which is a separate question from HIPAA compliance. Each clinical application would also need to be tested on its own task and population. These results tell us how Jev did on a synthetic benchmark; clinical diagnostic accuracy and the meaning of its probabilities for actual patients would need their own study.

Setting Up the Test

Like with a lot of my projects these days, I started out in ChatGPT. After a good amount of searching, I looked through a few papers and public datasets and chose MentalBench. Codex did essentially all the coding and assisted with the analysis and editing. We talked over different interpretations of the data, although the final interpretations are my own. We talked about how much instruction to give in the prompt and ultimately decided on using information from DSM-5 criteria for the 23 diagnoses, which span quite a range of different types of diagnoses. When doing some of the initial testing for accuracy, I saw difficulty differentiating certain diagnoses like OCD, ADHD, and phobias. So we beefed up those diagnostic definitions in general to help Jev decide when a diagnosis applied.

The way the researchers created this synthetic data is that they had different clinical seeds and created multiple versions of cases from those seeds. We didn't want related cases to overlap between what we explored and what was held out, so we grouped them using case identifiers and option patterns before splitting the data. After keeping things clean that way, we ended up with 1,965 cases, or 7.94%, to explore and 22,785, or 92.06%, held out for the final test. I also told Codex never to look at the holdout case text during development. Neither I nor the AI looked at it while we refined the prompt and criteria.

We decided to give Jev a yes-or-no option for each available diagnosis as part of the prompt, all in a single API call per case. That way we got a probability per diagnosis and could look at different thresholds for when a diagnosis counted as included. That was a decision I made after our initial exploration of Jev on the development data. For Types 1 and 2, which have one answer, we took the top diagnosis. For Types 3 and 4, we included every offered diagnosis scoring at least 0.50 and required that set to match the benchmark answer exactly. We used 0.50 because it gave the highest paper-weighted exact-match score among the cutoffs we tried on the exploration cases. We also looked at what would happen with 0.40 as a more inclusive Type 3 comparison. Then we froze the instructions, diagnostic guide, model version (jev-1.13.0), and scoring rule before running the holdout. These cutoffs are rules for scoring this benchmark, not clinical thresholds.

What Jev Did

Jev's paper-weighted exact-match score on our four-option holdout was 62.8%, with a 95% interval of 61.4% to 64.2% based on the reconstructed case groups. The paper's best published overall score was 62.69% for Claude Sonnet 4.5. So Jev was in about the same range. Our run used a held-out subset and our DSM-informed guide, while the paper reported results for the full dataset under its own instructions. We weighted each type by its share of the full dataset, as the paper did. Type 4 accounts for 13,050 of the 24,750 cases, about 53% of the overall score. That's why even with high Type 1 and 2 results, the overall score was about 63%.

Case typeJev, four options on our holdoutClaude Sonnet 4.5 in the paper
Type 194.6%94.32%
Type 288.9%91.16%
Type 332.4%14.99%
Type 466.9%74.84%
Paper-weighted overall62.8%62.69%

Jev did great on Types 1 and 2. I expected that because they ask for a single diagnosis and are more clear-cut. Type 2, told from the patient's point of view, also went very well. To me, that shows Jev finding structure within a more narrative, unstructured style, which speaks to its ability as a fuzzy classifier. Type 3 was more difficult. Jev's top answer matched at least one of the two accepted diagnoses in 95.5% of these cases, but at our 0.50 cutoff it returned the exact pair in 32.4%. Lowering the cutoff to the previously selected 0.40 point raised exact-pair agreement to 40.2%, while making the answer set more inclusive. The 32.4% was above the published Claude Sonnet 4.5 and GPT-5.1 figures, although Qwen 3 235B reached 54.19%. In my mind, Type 3 is where clinical reasoning is really involved, especially with information missing. Jev isn't built for that kind of deliberate reasoning. But with the right amount of structure, it was able to see through some of the fog in these cases. I wonder whether a different prompt could improve Type 3 accuracy, or whether more deliberate reasoning is needed to weigh several possibilities at once. I didn't spend this project tuning a separate Type 3 prompt. We used the same question wording and diagnostic guide across all four case types, which also let us see how far one general setup could go.

Type 4 is more about which one of two close diagnoses fits. Jev's top answer was the benchmark diagnosis in 88.0% of these cases. But its exact-set score was 66.9%, because an additional diagnosis above the cutoff also made the answer wrong under the paper's scoring rule. I thought this was a good result, even though it pulled down the overall score because Type 4 makes up so much of the benchmark.

I also wanted to see what would happen when Jev could choose from all 23 diagnoses. The table below shows how often its top diagnosis matched one of the benchmark's keyed answers. For Type 3, this top-answer measure credits one member of the pair.

Case typeAll 23 available, top answer matched a benchmark key
Type 191.9%
Type 274.1%
Type 374.9%
Type 477.5%

Using a version of the paper's weighted scoring rule produces an all-23 proxy of 58.4%. This scoring looks only at the benchmark labels among the four offered options, while Jev assessed all 23. The benchmark leaves diagnoses outside the offered four unscored. Giving Jev four options helped it find the benchmark answer more often, while the all-23 test asked it to distinguish among a much broader set of diagnoses on every case.

We made 49,500 successful Jev requests across the full exploration and holdout data, one request per case in each condition. Estimated API cost was $2.95 for the four-option condition and $13.77 for the all-23 condition, about $16.72 combined, excluding the small pilots and smoke test. On the holdout, median request times were about 0.22 seconds with four options and 0.32 seconds with all 23. Both holdout conditions together took 70.6 minutes of active run time with three workers. I was able to set up and complete this project in one day through collaboration with Codex. The project repository includes the protocol, code, reports, and complete results bundle.

The MentalBench paper reports accuracy, but we don't know what the models cost to run or how long they took. We can make a rough comparison using Anthropic's published API prices for Claude Sonnet 4.5. If its input were about the same as our four-option Jev input total of 70.2 million tokens, and it produced 20 output tokens per case, that would be about $218 at standard rates or $109 through its Batch API, before prompt-cache savings. Claude could tokenize the cases differently, and its actual prompts, outputs, and caching are unknown, so these are illustrative estimates, not the paper's measured costs.

Imagining Possibilities With System One Models

What might System One models like Jev make possible? We don't know how progress in this area will go, but TypeSafe calls Jev its first public model, currently in early access, and says it's still in the early days. What could assessment look like when classification can happen at the speed of discussion?

That starts to look more like assessment as an ongoing process, rather than just a single form or one-time diagnosis. Throughout an intake interview, information could be gathered as it comes in, and the clinical interview could adapt over time. A clinician could use that information in the moment and then do further analysis before making a final judgment. Jev took text in this project. I wonder whether audio or visual cues could be translated into text or structured code that a classifier could act on. That would be another workflow to build and test, but it's an interesting possibility for assessment. It's a sort of measurement-based approach that would need validation.

Coming back to the idea of making unstructured data usable for research, I think today a lot of psychological data is self-report. Often it's in a measurement form; sometimes it's a transcript, and there's a whole range between more and less structured. Being able to ask new questions of that material could help bring qualitative and quantitative research together. I also think about wearable and smartwatch data already being captured and what we might do with it if it could be represented in a form a classifier can use.

For quality improvement, imagine asking a question about medical records that a particular EHR had never explicitly tracked. The information may already be in notes or other unstructured elements. You could ask a System One model whether the records meet certain qualifications across thousands of instances, gathering answers from work clinicians are already doing. They wouldn't have to adopt a new workflow, press new buttons, or take another survey just so we can ask that question.

I can imagine an intake system, human-led or AI-led or both, helping route people based on their presentation, using a System One model for real-time decisions. What level of care should they go to? Should they go to group or individual therapy? What's the level of safety concern? What needs to be elevated or not? Suggestions could go to the clinician or person making those decisions. I also imagine an ongoing conversation, a chart, or a chatbot with several classifiers looking at things like safety, context, or treatment decision points. Different probability cutoffs could trigger different reactions or warnings. The system could flag the right person, remind a clinician of a protocol, or give an LLM options for how to respond next. These little points of decision-making could be built into the EHR, note-taking software, or other software psychologists use.

The way I see it, what might have required machine learning engineers and the creation or fine-tuning of a model to predict, group, or measure data can start with questions posed by domain experts. In psychology, that might be the psychologist or researcher describing the task in natural language. You still need validation data and need to know what makes the model useful for a given task in a given setting. But the barrier to trying it out seems much lower. You can change the qualifications, test again, and, if the performance is good enough, use it at scale affordably.

I see that as a kind of democratization of classification models. It's sort of like taking vibe coding into vibe classifying. Other grad students like myself, and people working in academia who may have limited funding, could benefit from affordable and fast models like this. You can work with datasets at scale and potentially find insights that are more generalizable, depending on the data and who it represents. As long as there's ethical access to a dataset for that use, the speed and cost can open up more research opportunities. And with more creative use in medical records or therapy sessions, System One models could work behind the scenes on decision-making, observation, or classification tasks. I think we can make a lot more use of the psychological data we have, and find creative and meaningful uses for technology like this in psychological science and practice.

Process note. I supplied the questions, chose the dataset, developed the clinical framing, and talked through the interpretations and ideas in this post. ChatGPT helped me find candidate datasets. Codex supported the code, research, analysis, and editing. I also used Codex to revise this article against my original spoken riffs.

Subscribe for future posts

If you want new writing at the intersection of AI and psychology, ethics, and implementation of AI in clinical practice, subscribe on Substack.

Subscribe on Substack

The views expressed here are my own and do not necessarily reflect the views of any current or future employer, training site, academic institution, or affiliated organization.