Self Research Institute

@selfresearch.org

Nonprofit building open, self-sovereign tools for personal health data. Personal informatics · whole health · open science. Your data, owned by you. → selfresearch.org

A database schema tells you what columns exist. An ontology tells you what things are—and how they relate. Heart [is-a] organ. Organ [part-of] cardiovascular system. These relationships aren't just metadata—they ARE the model. 🧵👇

An infographic about Data Modeling on a light grid background. The headline reads: "A schema tells you what columns exist. An ontology tells you what things are."

A "Schema" section shows flat database columns: patient_id, organ, and diagnosis_code. Below it, an "Ontology" section illustrates these concepts as a connected knowledge graph: 'Heart' has an 'is-a' relationship to 'Organ'; 'Organ' has a 'part-of' relationship to 'Cardiovascular system'; and 'Infarction' has an 'affects' relationship to 'Heart'.

Text below reads: "Those relationships aren't metadata sitting beside the data. They are the model — which is how an ontology answers questions a spreadsheet structurally can't."

A dark blue banner at the bottom states in bold white text: "WE BUILD ONTOCODE FOR PEOPLE WHO MODEL MEANING, NOT JUST STORE IT." The footer includes the Self Research Institute logo and website url.

A dataset without metadata is a locked box—technically there, but practically useless. Metadata (what, how, when) is the single biggest factor in whether research data gets reused or quietly expires. It's the difference between data with a future and data with none. 📦🗝️ #Metadata #FAIRData

An infographic by Self Research Institute on the topic of Data Stewardship. The main headline reads, "A dataset without metadata is a locked box. Technically there, but practically useless to anyone except the person who made it." Below is a section titled "The description that unlocks it," showing an example of metadata: "What was measured" is "Resting heart rate, bpm"; "How" is "Chest-strap ECG, 1 Hz"; and "When" is "2025-03 to 2025-09". Below this, text reads, "It's the single biggest factor in whether research data ever gets reused instead of quietly expiring at the end of a study." The graphic concludes with a dark blue banner stating: "UNGLAMOROUS. ALSO THE DIFFERENCE BETWEEN DATA WITH A FUTURE AND DATA WITH NONE.”

Normally, AI training means gathering massive data in one place. Federated learning flips that: the model travels to the data. It trains locally, and only learned patterns—never raw datasets—are shared back. Institutions keep their data secure while AI still learns. 🧠🛡️ #FederatedLearning

An infographic by Self Research Institute explaining Federated Learning. The main heading reads, "Instead of moving the data, the model travels to the data." A diagram shows three separate locations: "Hospital A," "Registry B," and "Lab C." Inside each location, text notes that "Raw data stays inside" and the model "Trains locally." Arrows indicate that the "Model goes out" to these locations, and "Patterns come back" to a "Central model," which "Receives only the learned patterns." The graphic concludes with the text: "AI can learn from many institutions' worth of data, without any one of them handing over proprietary or sensitive information. THE PATTERNS TRAVEL. THE DATA DOESN'T.”

Since 2021, U.S. providers must legally give patients fast, free electronic access to their health data under the 21st Century Cures Act. Deliberately or lazily blocking access carries real penalties. Policy finally caught up to interoperability. Details & sources 👇

"Infographic by Self Research Institute about the Information Blocking Rule. The main text explains that since 2021, U.S. healthcare providers have been legally required to give patients fast, free electronic access to their own health information. It highlights that blocking this access, deliberately or through inaction, can carry real penalties under the 21st Century Cures Act. A concluding point notes that the law's premise is that technical ability isn't enough if incentives point the other way, meaning policy caught up to the interoperability problem before most infrastructure did. The footer includes the Self Research Institute logo and website URL.”

FHIR gets used as shorthand for "health data interoperability, solved"—but it isn't quite that. FHIR is a standard for exchanging healthcare data using modern web technology (the exact same REST APIs & JSON that power most apps you use daily). 🧵👇

"Infographic by Self Research Institute titled 'HL7 FHIR • WHAT IT IS, WHAT IT ISN'T'. The main headline reads, 'FHIR is the delivery truck. Not the label on the cargo.' The graphic is split into two columns. On the left, a section labeled 'WHAT IT DOES' explains that FHIR 'Moves healthcare data over the same REST APIs and JSON that power everyday apps — split into modular resources.' It lists examples like 'Patient', 'Observation', and 'Medication', alongside a small JSON code snippet. On the right, a section labeled 'WHAT IT DOESN'T' states it does not 'Guarantee that two systems mean the same thing by those resources. A shared envelope isn't shared meaning.' It illustrates this with 'Observation =? Observation' and includes a box stating, 'That still takes terminologies and ontologies.' A thick dark blue banner at the bottom emphasizes in bold, 'TRANSPORT SOLVED. MEANING STILL NEEDS A SHARED LABEL.' The footer contains the Self Research Institute logo and website URL.”

Most disease outbreaks don’t start with humans—they start in animals, water, or stressed ecosystems. One Health treats human, animal & eco health as one system. The hard part isn't the idea. It's building data infrastructure so labs in different sectors can read each other's data 📊🧬

"Infographic by Self Research Institute titled 'ONE HEALTH'. The main headline reads, 'Three sectors. One interconnected system.' Below it is the subtext: 'Most emerging human infections begin somewhere other than a human — an animal, a water source, an ecosystem under stress.' A triangular diagram in the center illustrates 'One Health - TREATED AS ONE', showing bidirectional arrows connecting three nodes: 'Human health' at the top, 'Animal health' at the bottom left, and 'Environment' at the bottom right. A banner below labeled 'THE HARD PART' states: 'Not the idea — the shared data infrastructure that lets one sector actually read another’s data.' A thick dark blue banner at the bottom emphasizes in bold, 'HEALTH DOESN'T RESPECT SECTOR BOUNDARIES.' The footer contains the Self Research Institute logo and website URL.”

A digital twin in healthcare is a living virtual model of a patient—built from genetics, wearables & medical history. Clinicians use it to simulate how a disease progresses or a treatment responds before trying it on you. It’s the end of "one-size-fits-all" medicine. #HealthTech

Infographic by Self Research Institute titled 'DIGITAL TWINS • ACTIVE RESEARCH, NOT SCIENCE FICTION'. The main headline reads, 'A virtual you, continuously updated.' The graphic illustrates the data flow of a digital twin. On the left, under 'BUILT FROM', are four inputs: Genetics, Wearables, Imaging, and Medical history. Arrows from these inputs point to a large teal box on the right labeled 'YOUR DIGITAL TWIN: A living model of one person'. Below this, under 'USED TO SIMULATE', are two boxes: 'How a disease might progress' and 'How a treatment might respond', followed by the caption '— before anything is tried on the real person.' A bright purple banner at the bottom states in bold, 'MEDICINE BUILT AROUND ONE PERSON'S ACTUAL DATA.' The footer includes the Self Research Institute logo and website URL.”

Who actually owns your health data? Patients assume it's theirs, biobanks claim rights, and genome projects call it public. But "ownership" is the wrong frame entirely. What actually matters is who controls access, under what conditions, and for how long. That's the real fight.

Infographic by Self Research Institute titled 'HEALTH DATA SOVEREIGNTY'. The main headline reads, 'Three parties. One dataset. All claim it.' Below are three white cards showing different perspectives on data ownership: 'NATIONAL GENOME PROJECTS' claiming 'Publicly owned.', 'BIOBANKS' stating 'We hold rights to coded data.', and 'PATIENTS' saying 'Obviously, it’s mine.' A pale red warning banner below states, '"Ownership" is often the wrong frame entirely.' A navy blue panel titled 'WHAT DECIDES IT IN PRACTICE' lists three numbered points: '01 Who controls access to it?', '02 Under what conditions?', and '03 For how long?'. A final bold dark blue banner at the bottom concludes, 'CONTROL, NOT OWNERSHIP, IS THE REAL FIGHT.' The footer includes the Self Research Institute logo and website URL.”

Findable. Accessible. Interoperable. Reusable. Since 2016, the FAIR principles have shaped health data management because so much valuable data is collected once & never used again. FAIR isn't a buzzword; it's a checklist to ensure data outlives the study that produced it.

"Infographic by Self Research Institute titled 'FAIR DATA • SINCE 2016'. The main headline reads, 'Four words that quietly shape modern health research.' accompanied by the subtext, 'They exist because so much valuable research data is collected once — then never usable again.' Below are four cards explaining the FAIR acronym: 'F - Findable: Described well enough that someone can locate it at all.', 'A - Accessible: Retrievable under clear, stated conditions — open or controlled.', 'I - Interoperable: Uses shared vocabularies, so other systems understand it.', and 'R - Reusable: Documented and licensed well enough to use again, elsewhere.' A small banner below states, 'Not a certification. Not a buzzword. A practical checklist.' A large dark blue banner at the bottom concludes with, 'DATA SHOULD OUTLIVE THE STUDY THAT PRODUCED IT.' The footer includes the Self Research Institute logo and website URL.”

Fragmented health data isn't just an inconvenience—it's billions in unnecessary tests, delays & admin overhead. Patients pay thousands more out of pocket. The tech to fix this exists. The bottleneck isn't computing power—it's that systems don't share a common way to describe what data means.

"Infographic by Self Research Institute titled 'THE COST OF FRAGMENTED HEALTH DATA'. The main headline reads, 'Not an inconvenience. Billions a year.' Below are three white cards detailing these costs: 1) 'Duplicate tests: Imaging and labs re-run because prior results can't be found or read.' 2) 'Delayed treatment: Care waits while records are chased across unconnected systems.' 3) 'Admin overhead: Staff hours spent reconciling records by hand instead of treating patients.' Beneath the cards, an outlined banner states: 'And patients in highly fragmented care pay thousands more out of pocket each year than those whose care is connected.' A dark blue section follows, contrasting 'NOT THE BOTTLENECK' with the words 'Computing power' crossed out, against 'THE ACTUAL BOTTLENECK: No shared way of describing what the data means.' A final bold banner at the bottom reads, 'THE TECHNOLOGY TO FIX THIS ALREADY EXISTS.' The footer includes the Self Research Institute logo and website URL.”

One hospital logs "BP", another "blood pressure", another "systolic/diastolic". To a human, it’s obvious. To a computer, it’s 3 unrelated words. An ontology fixes this by providing formal, shared definitions. It's the difference between data that’s merely stored and data that’s understood.

An infographic by Self Research Institute titled 'WHAT AN ONTOLOGY ACTUALLY FIXES'. The graphic contrasts how humans understand data versus how machines see it. The main headline states: 'Obvious to you. Three unrelated words to a machine.' Below the headline, three separate sources are shown using different labels: Hospital A uses 'BP', a Health App uses 'blood pressure', and a Lab Report uses 'systolic/diastolic'. A broken gray line connects them to an orange warning banner that states: 'No shared definition — the machine sees three strings, not one concept.' The graphic then presents the solution: a navy blue 'ONTOLOGY' box that provides 'ONE CLINICAL CONCEPT' named 'Blood pressure'. Converging arrows illustrate mapping the three different inputs into this single, central definition. Tags characterize the ontology concept as having a 'formal definition', being 'shared across systems', and being 'machine-resolvable'. A final text banner at the bottom summarizes the core value: 'The difference between data that’s stored and data that’s actually understood.' The footer includes 'SELF RESEARCH INSTITUTE | WWW.SELFRESEARCH.ORG' and the SRI logo.

💡 The fix for LLM hallucination isn’t a better prompt. It’s structure. Standard RAG forces an LLM to hunt through unstructured text, leading to lost context. A Semantic Layer feeds the AI a structured knowledge graph, delivering grounded, reliable outputs.

Infographic titled "The fix for hallucination isn’t a better prompt. It’s structure." comparing Standard RAG vs. a Semantic Layer.

Left side (Standard RAG): Shows an LLM searching a messy, unstructured pile of text documents, resulting in "Hallucinations & lost context."

Right side (Semantic Layer): Shows an LLM reading a clean, structured node-link knowledge graph, resulting in "Grounded & reliable output."

A banner at the bottom announces: "Coming soon: our VS Code suite brings native semantic design right into your IDE — no jumping to separate modeling software." Developed by Self Research Institute.

The gap between academia and modern tech is really just a gap between file formats. 📄➡️💻 Bio-curation leans on strict OWL. Web devs need JSON-LD. For years, bridging the two has meant maintaining brittle, error-prone custom conversion scripts between your ontology and your product. (1/3)

A promotional graphic by the Self Research Institute for OntoCode. The top tag reads, "ONTOCODE • ONE SOURCE, EVERY FORMAT." The main bold headline states, "Curate in OWL. Export to JSON-LD." The subtext explains, "Academia leans on strict OWL; web builders need JSON-LD. Curate once, then export to the serialization your pipeline expects — no manual conversion." The central graphic is a dark-mode user interface showing a "File > Export As..." menu. On the left, a source file named "SDA_Merged.ttl" is shown with the note, "One canonical model. Edit it once." An arrow points to a list of seven export formats on the right: .owl (RDF/XML) tagged for "LAB", .jsonld (JSON-LD) tagged for "WEB" and highlighted, .ttl (Turtle), .owlxml (OWL/XML), .omn (Manchester), .ofn (Functional), and .obo (OBO Flatfile). The bottom status bar reads "one source of truth | 7 export formats | no manual conversion." A teal call-to-action button at the bottom says, "ACCESS THE BETA → Export in the format that fits your pipeline — link in bio.”

Standard JSON will break your N=1 health baseline. 📉 If you try to build individualized health models using flat JSON, your data architecture will eventually collapse under its own weight. Here is why strict JSON-LD serialization is a hard requirement for the future of precision health: 👇 (1/3)

An infographic by the Self Research Institute titled "N=1 Baselines • Why JSON-LD is Non-Negotiable." The main headline reads, "Three APIs. One heartbeat. Flat JSON sees three metrics." On the left, three dark boxes show different JSON outputs for a heart rate of 62: a Wearable API using "hr", a Hospital portal using "HeartRate", and a Smart ring using "pulse_bpm". Arrows point from these three boxes through a tag labeled "@context" and into a single box on the right titled "One Ontological Concept." This box displays the concept ":HeartRate" and explains, "JSON-LD maps every vendor's key to the same concept — so your baseline stays one graph, not three silos." At the bottom are two additional feature boxes. The first is "Time-series alignment: Link 1Hz wearable streams and once-a-year labs to one temporal ontology." The second is "Bridges OWL & the web: Ontologists keep strict OWL; engineers get the JSON their APIs expect.”

Stop copy-pasting citations into your workspace. 📄➡️💻 Sci2Code links your Zotero libraries directly into VS Code, pulling literature references straight into your code comments (Python, JS, R, Julia). The repo is officially public under AGPLv3. Star it here: github.com/The-Self-Res...

GitHub - The-Self-Research-Institute/Sci2Code-extension-for-vscode

Contribute to The-Self-Research-Institute/Sci2Code-extension-for-vscode development by creating an account on GitHub.

github.com

Beta registrations are open! 🕸️ We’re offering two ways to use our ontology editor, with one catch: 🌐 Browser Web App (Capped Seats): Runs on our servers. We’re capping seats to guarantee a fast, stable experience. Once full, they're full. 🧵

An infographic by the Self Research Institute for the OntoCode Beta titled "Pick your seat before the weekend." The subtitle reads: "Web seats are capped to keep the browser version fast and stable. Desktop IDE seats are open — no cap." Below are two side-by-side option cards. The left card, "Browser version," features badges for "Limited" and "Capped Pool," explaining: "A fixed number of web seats, so we can watch server load and keep it snappy." The right card, "Desktop IDE," features badges for "Open" and "No Limit • Install Today," explaining: "Runs on your own machine, so seats stay wide open. Reserve and download anytime." A dark banner at the bottom reads "Reserve your version before the weekend" alongside a "Reserve Now" button.

We map the stars, the oceans, and our streets. 🗺️ But when it comes to the systemic cascades inside our own bodies, we rely on fragmented PDFs and siloed dashboards. The Self Data Atlas is the architectural shift required to fix this. 🧵

Bild

The hardest problem in health tech? Time-series alignment. ⏱️ Harmonizing a 1Hz wearable stream with an annual lipid panel is a nightmare. The Self Data Atlas solves this using temporal ontologies—standardizing time so you can query mismatched data together. How do you align yours?

An infographic by the Self Research Institute titled "It's not the volume. It's the timing." It explains how the Self Data Atlas uses temporal ontologies to align different health data clocks. Below the text is a dark UI dashboard showing three separate data streams aligned on a single grid: a continuous blue wave for "Heart Rate – 1 Hz," a yellow dotted line graph for "Mood Score – 1x / day," and a single red dot for "Lipid Panel – 1x / year." The bottom text reads: "Same grid. Three granularities. One queryable timeline.”