Coenfirmation Bias

NLEN

The ScienceChecker app

Someone writes that coffee is healthy. Or that sugar is better for your muscles than soy. There is a source with it, a real study. So it is true. Right?

Not necessarily. A source under a claim only says that there is a source under it. Whether that study really supports the claim is another question. I ask that question in my blogs, my podcast and my workshops. Now I am turning it into an app. This is the story of that app: where it started, what went wrong along the way, and where it stands now.

Where it started

It did not start with an app, but with a form. To check a claim, I had a fixed scheme. Which source is it? What question did the researchers ask? What did the study look like? What did they find? And does that fit what the claim says?

Screenshot of the help page of the ScienceChecker. Under the heading 'What does the ScienceChecker do?' it says that the app does not judge whether a claim is true, but whether one particular study backs up that claim. 'Not supported' means that this study does not back up the claim. 'Cannot assess' means that the app is missing something it needs to make a judgement, for example because there is only a summary or because the claim is too vague.

Besides that, there is the method I teach in workshops: sciencechecking in three steps.

The question. Did the study actually investigate what the claim claims?

The comparison. Was the comparison made in a way that fits the claim? In nutrition, it is not about good or bad, but about better or worse than something else.

The answer. Does the outcome of the study say what the claim says?

Sciencechecking is not fact-checking. The question is not whether a claim is true. The question is whether this source supports this claim. A claim can be right and still be poorly supported. Then the app says: not supported. That is not a mistake. That is exactly the point.

The mission

Screenshot of the input screen of the ScienceChecker. Step 01 is the claim, with the example 'Omega-3 lowers LDL cholesterol', and a box for people who do not have a claim yet and want to see what a study can support. Step 02 is the source: paste a link, DOI or PubMed ID, or upload a pdf. At the bottom is the button 'Start check'.

I want you not to simply believe sources, but to be able to check them. Most people skip the sources. That makes sense: reading a study takes time and requires knowledge. The app has to make that easier. You enter a claim and a study. The app goes through the three steps and shows where it holds up and where it does not.

My motto applies here too: it is not about what the truth is, but about how you arrive at the truth. So the app does not only give a judgement. It shows the reasoning, step by step, so that you can follow it and check it yourself.

The steps so far

A first version in a few hours. I built the app with Emergent. That is a program in which you describe in plain language what you want, after which an AI builds the app for you. I have no experience building apps. Still, within a few hours there was something that worked: claim in, study in, judgement out. I thought the work was almost done. It had only just begun.

Writing down the method. When I assess a study myself, I do many things without thinking about them. An app cannot do that. Everything has to be written down, as a rule, with examples. So the method soon went through a series of versions. Some things were also dropped. The first plan had the app also weigh the total evidence on a subject. It no longer does that. It assesses one source for one claim, and nothing else.

The first testers, in June. A small group, slightly fewer than ten people, did 35 checks together. The reactions were positive. But I also saw that people without experience had trouble with the app. From that I made a list of improvements. Some of them are in the app now. You can tap difficult words for an explanation. There is a help page. There is a copy button with a short answer that you can send back to whoever made the claim. And if you enter only a study, the app itself formulates the strongest claim you can base on that study. Other points, such as an account page of your own and a place to ask questions, are still waiting. In addition, the app now works in Dutch and in English.

Top of the ScienceChecker, with a 'Help' button, a choice between NL and EN, and the subtitle 'Test nutrition and health claims against the actual study.'

August: the first measurement. At the start of August, the three errors that were then still on the list for the launch were fixed. After that, the work was mostly about the method itself. In mid-August came the first measurement with a fixed test set. More on that below.

September: building and measuring. In September, a new series of five building rounds followed. After each round, it is checked whether the change is really in the code, and not only in the report. A second, separate session reviewed the last round and found new points. Those are now on the list.

Why the method is so much work

Two examples show why writing down the method takes so much time.

Someone claims something about muscle building. The study measured fat-free mass, so the weight of everything except fat. Is that the same? Not quite. But is it then immediately wrong? No, that is not so either. The answer became: it may go on to the next step, but the app must name the difference. And it must name the difference before the figures, not after. Otherwise you read the figures as the answer to a question that was not asked.

Or take "Coffee is healthy". Sounds like a claim, but what is healthy? You cannot measure that. The app has to stop here and say that the claim is too vague to test. "Coffee lowers blood pressure" it can test. The difference between those two sentences had to be fixed in a rule.

The first full version of the method was version 3. Now it is version 6.4. Everything is in three documents: one about how a source is assessed, one about the technology and the costs, and one about what the user sees on the screen. There is also a log in which every decision is written down, with the reason.

The biggest challenges

A language model does not always give the same answer. Behind the app is a language model: an AI that reads and writes text. In my case that is Claude, from Anthropic. If you ask it the same question twice, you do not always get exactly the same answer. In the first measurement, the final judgement within each test case was always the same. But the details around it varied: the remarks with a judgement and the assessment per step. For an app that judges science, those details have to be right too.

The model follows an example sooner than a rule. That was a surprise. In an earlier version, a rule and an example did not say exactly the same thing. I saw that clash eight times in the answers of the model. All eight times the model chose the reading of the example, not that of the rule. In this case, then, the example weighed more than the rule. That is why every example has to match the rule above it exactly.

Measuring instead of believing. That is why a fixed test set was made: a series of claims with a source, for which the right judgement was fixed in advance. There are now sixteen cases in it. Each case is supposed to run several times, and the outcome is laid next to the expected judgement. In the last measurement, fifteen of those sixteen cases were tested. Of those, thirteen came out right on the judgement. That sounds good, but it does not say much yet. Most of the cases were run only once, and the measurement was from before the last building round. And thirteen out of fifteen is not enough in any case. The three steps and the final judgement have to be one hundred per cent right. Not nearly. One hundred.

My own record-keeping. Anyone who works like this writes a lot down. And what is written in two places will sooner or later differ. That is why every fact has one place of its own. Small checking programs look at whether the documents still belong together and do not contradict each other. And I work with a fixed set of working rules. There are 39 of them now. Each rule comes from a mistake that was really made.

How I work

I do not build the app alone. Emergent carries out the building. Claude, an AI assistant, writes the building instructions, lately with the code included, checks the result and keeps the documents up to date. That is the same name as the language model in the app, but a different role. In the app, Claude assesses a source. Here, Claude helps me build. I decide about the method. Claude puts the choice to me, with the strongest objection included. Smaller choices I sometimes leave to Claude, and I check those afterwards.

That does not mean that I trust the AI. The opposite. A second, separate session reads through the work of the first. It gets the files themselves, not a summary, and looks for mistakes. That has already turned up mistakes many times that would otherwise have stayed in. So what I do with studies, I also do with the AI that I use myself. Whoever claims something has to show it. Also an AI. Also me.

Where we stand now

The app runs, but it is not public yet. The fifth building round is finished. The measurement that has to show whether that round improved anything is ready.

Screenshot of the history in the ScienceChecker, with 50 checks. Four earlier checks are in view, each with a judgement: 'Supported' for a claim about whey protein and fat-free mass, 'Not supported' for vitamin D and cancer, 'Cannot assess' for omega-3 and LDL cholesterol with a PubMed link as the source, and 'Not supported' for a claim about the absorption of animal and plant protein.

Before the app can go public, a few things still have to happen. Among other things security and accounts, and agreements about costs and use. The condition that weighs heaviest for me: the test set has to show that the app gets the three steps and the final judgement right without errors. Even then, something is left to check. The test set tests the judgement of the model, not everything the user sees on the screen.