Skip to content

Reviewing a DisMech Disease Report — Instructions for Clinical Reviewers

What this is

DisMech is an AI curation system for disease pathophysiology. We are trying to understand how far we can get with unsupervised, AI-driven curation — how much of a rigorous, referenced disease record an AI can build with no human in the loop, and, crucially, where and how it fails. More about the project: dismech.monarchinitiative.org (code at github.com/monarch-initiative/dismech).

The document you are about to review was written entirely by that system. Your review is how we find out whether it can be trusted.

Your role

Read the report the way you would grade a bright but unsupervised junior colleague — a student or new staff member who has handed you a draft — or the way you would review a manuscript submitted to a journal. Assume nothing is correct just because it is confidently written. The system writes fluently even when it is wrong; polished prose is not evidence.

You do not need to fix anything or rewrite text. Your job is to judge it and flag what is wrong.

How to do it

  1. Open the report in Google Docs (or Word) — you'll be given a link.
  2. Read it top to bottom.
  3. Leave your reactions as comments on the relevant text. Highlight the span, add a comment. That's it. Short is fine.
  4. For a reaction that isn't tied to one passage — an overall impression, a cross-cutting concern — add it as a comment at the very top of the document.

Lean on your own clinical judgement first. You do not need to read the cited sources to review the report, though you are welcome to open one if a particular claim looks off. And when you disagree, giving a citation of your own is welcome but never required — a blunt "WRONG — X is actually Y" is exactly as useful to us as a referenced one.

Each factual claim in the report comes with its supporting reference, an exact quoted snippet from that source, and an interpretation. That reference is where the AI found the information it then drew conclusions from, so if something doesn't look right, it could have come from — or been misinterpreted from — that source. Feel free to point at it in your comment.

What we're looking for — in priority order

  1. Inaccuracies and outright falsehoods — the top priority. Anything that is wrong, overstated, out of date, or clinically implausible. Don't overthink the wording — a comment that just says "WRONG" plus a few words of why is exactly what we want. Examples:

WRONG — riluzole extends survival by ~3 months, it is not "curative".

Overstated — this gene is a rare modifier, not a primary cause.

This was the view 15 years ago; the current consensus is X.

The cited source is real, but the terminology/criteria have changed since it was published — or it conflicts with [guideline/source] and that conflict isn't noted.

  1. Gaps and omissions — nearly as important. Things a complete record should contain but this one doesn't. This is where your field knowledge is most valuable, because the AI can't miss what it never saw. Examples:

Modern practice recognizes ~12 subtypes; only 6 are listed. Missing: …

C9orf72 is the single most common genetic cause and isn't mentioned.

No mention of [key treatment / biomarker / diagnostic criterion].

  1. Structural / presentation notes — minor. Anything about how the report is organized, ordered, or laid out — sections in the wrong place, confusing framing, redundancy, missing headings you'd expect. A quick note is plenty.

  2. Things done well — a bonus, optional but genuinely useful. Mark anything that is especially correct, well-synthesized, subtle, or interesting. Knowing what the system already gets right is as informative for us as knowing where it breaks — and it helps us avoid "fixing" what isn't broken. A simple "good" or "nicely put" in the margin is enough.

What happens to your comments

We take every comment and map it to a specific AI failure (or success) mode — for example:

  • Did the AI misinterpret a phrase or finding in a paper?
  • Did it extract part of the picture and then get lazy — capturing some aspects a source mentions but dropping others?
  • Was the source itself bad — outdated, low quality, or wrong for this claim?
  • Did it hallucinate content, a citation, or a detail outright?
  • Or did it get something right in a way we should reinforce?

You don't need to classify anything yourself — that's our job. Just react honestly and we'll do the mapping.

Anything else is welcome

You are among the first people to review AI-generated disease records at this depth. If, along the way, you form an opinion about how these documents should be reviewed — a better workflow, a section that deserves more scrutiny, a type of error we should be systematically watching for, a way the report could make your job easier — please tell us. A note at the top of the document, or a reply to whoever sent you the report, is perfect. That meta-feedback may be the most valuable thing you give us.

Credit

All reviewers will be offered co-authorship on a potential evaluation paper about this work. We can't promise that paper will happen, but if it does, you'll be invited onto it (should you want to be).

Thank you. Be blunt — bluntness is the point.