The dataset behind our topic choices
BeatSIBO decides what to write about by looking at what patients actually discuss, rather than at what a keyword tool suggests. This page explains what that means, and what it does not mean.
What it is
The short answer
695,050 posts and 7,205,554 comments from 18 patient communities, covering 2011-02-24 to 2026-05-26. It is a database read offline only. It is never queried when you load a page, it is not deployed with this site, and no part of it is republished here.The communities cover SIBO, IBS, mast cell activation, histamine intolerance, long COVID, dysautonomia, mold illness, candida and food allergy. The dataset itself is public: it is published at huggingface.co/datasets/toczix/sibo-research-db and github.com/toczix/sibo-research-db, with SHA-256 reddit.db@a313e531.
What it decides
One thing: which pages get written. We count how often two conditions are discussed in the same comment, and we write a page for a pair only where that count clears 2,000. The bar was set after seeing the whole distribution rather than before, because a threshold chosen in advance tends to admit exactly the pages you already wanted to write. Below the bar we would be guessing at what confuses people, and we would rather leave a gap than fill it with an assumption.
These counts do not appear on the pages themselves. They are how we choose the subject, not evidence about your body, and putting a comment count in front of someone trying to work out what is wrong with them confuses those two things.
What this data cannot tell you
This is the section most sites would leave out, and it is the reason to trust the rest of the page.
- It is not a prevalence estimate. These are people who posted on Reddit, not a sample of anyone. A count of comments tells you what a self-selected group discussed, and nothing whatsoever about how common a condition is.
- Nothing here is verified. No diagnosis in this dataset was confirmed. People report what they believe they have, which is frequently not what they have.
- Success-story communities are survivorship bias by design. A community that exists for recovery narratives contains recovery narratives. The people for whom nothing worked are, by construction, somewhere else.
- Co-mention is not correlation. Two conditions named in the same comment may be compared, confused, or merely listed. We use these counts to decide what to write about. We do not use them as evidence that two conditions are related, and neither should you.
- English-language and Reddit-demographic skewed. Younger, more online, more Western, and more likely to have already been dismissed by a doctor, which is a specific population with specific complaints.
- It is a snapshot. Coverage ends 2026-05-26. It does not update itself, and anything after that date is invisible to it.
What we will not do with it
We do not republish comment text at scale, we do not name or profile individual users, and we do not run a search front-end over it on this site. Aggregate analysis of public posts is ordinary. Republishing pseudonymous health disclosures on a page that earns money is a different act, and we are not going to do it.