From public data to usable evidence
How one failed health-system profile became a research architecture for source acquisition, identity, governed interpretation, and publication.
Role
Founder and operator
Open-Informatics 2026 to present
My contribution and results
My role in the work
I founded Open-Informatics and designed the research architecture described here. I developed FAST Identity as a completed research implementation, and I continue to build and maintain Healthcare Data MCP, Healthcare Agents, and the USHSO profile and evidence workflow. AJHCS is jointly built and retains independent scholarly decision-making; Open-Informatics designed and maintains its publishing infrastructure but does not control its editorial judgments.
Where it stands
The program now brings together working components for public-source acquisition, identity research, specialist workups, governed profiles, and open-access distribution. Their development has included repeated implementation tests, acceptance replays, adversarial cases, and performance experiments; the next phase is to turn that accumulated evidence into a public worked example, broaden source and state coverage, and complete the temporal and comparability layers.
Contents
The profile failed before the research was done
Health-systems research can look like a fact-gathering exercise, but the harder question is usually more basic: what, precisely, is the institution being described? A hospital, physician group, legal filer, facility roster, ownership relationship, and market-facing brand may all refer to different parts of the same enterprise, while their public records cover different periods and reporting perimeters.
The records are scattered across federal releases, state licensing systems, audited financial statements, Internal Revenue Service filings, institutional websites, and local reporting. AI has helped me find sources, extract candidate facts, compare approaches, and prepare early syntheses. The FAST development program added a systematic experimental record: acceptance replays, adversarial identity cases, exhaustive policy tests, live-source checks, measured rebuilds, and a benchmark harness with separated evaluation evidence. These are substantial first-party experiments across the research process.
One early development run made the limits concrete: reconciling bed counts for Jefferson Health and Lehigh Valley Health Network, only one section of a larger comparative profile, took more than 80 individual tool calls. Roughly 30 failed; the work consumed roughly 100,000 tokens and more than five minutes, and it still returned guessed counts while facility coverage, service capabilities, and outpatient inventories remained unreconciled. These approximate figures describe one first-party development run, not an independent benchmark. 1 The bed-count reconciliation The development record documents more than 80 calls, roughly 30 failures, roughly 100,000 tokens, more than five minutes, and an incomplete result. It is a design observation, not an independent benchmark. Open source Read full Note 1
AJHCS interns and research assistants were reaching the same boundary from different directions: access to a source was only the beginning, because a researcher still had to determine which entity the record described, preserve that distinction during a join, and produce an answer that another person could inspect. I created Open-Informatics to work on that boundary through a set of components that could keep claims bounded, retain source lineage, represent uncertainty, and leave release decisions with accountable people, rather than through another dashboard or a general assistant that could write past missing evidence.
An architecture developed through repeated testing
Open-Informatics is an implemented research architecture whose components have been developed through repeated testing at different levels of maturity; several contracts and repositories are public, FAST Identity has been exercised through acceptance and adversarial cases, and the wider program records failures, performance, and policy behavior as part of its development. The source-to-publication model is the program's organizing thesis, with a public worked example and a bounded evaluation of the integrated path forming the next stage of the research rather than a reason to discount the experiments already completed.
The institutions retain separate responsibilities across acquisition, identity, interpretation, and publication. Public access also varies by component.
Figure 1 · Conceptual model
A claim has to survive more than retrieval.
The architecture separates custody, identity, interpretation, comparison, and release. A missing or conflicting record is allowed to stop the path.
Read the figure as text
- Evidence path. A source artifact becomes a bounded observation, a resolved subject, a governed projection, and then a published claim. Normalization, identity adjudication, governance, and accountable release are separate gates. Missing, conflicting, stale, or out-of-scope evidence can lead to a hold, a question, or abstention.
- Comparability. Two records are eligible for comparison only when the institutional subject, observation period, unit and accounting basis, and included organizations align.
- Institutional loop. Open-Informatics develops methods, USHSO maintains reference records, and AJHCS turns evidence into scholarship. Questions and evidence gaps return to the research agenda. AJHCS keeps editorial authority, and software inherits neither source nor professional authority.
Research position
Established methods, applied to health-system research.
Open-Informatics draws on established work in provenance, entity resolution, knowledge graphs, temporal databases, and abstention, applying those traditions to the practical problem of building and publishing bounded claims about changing health systems from public records; the sources below locate the architecture within that research context without implying formal conformance.
Research objects need durable identifiers, reusable metadata, and a record of the entities, activities, and agents involved in their production.
Healthcare Data MCP carries source, vintage, producer, checksum, coverage, and conflict fields in its public evidence contract. The project does not claim formal FAIR or PROV-O conformance.
Entity resolution is a data-quality problem that joins descriptions of the same real-world object across heterogeneous sources.
FAST Identity applied that problem to health systems with fail-closed admission and separate legal, operational, and market-facing categories. It is a completed research implementation, not an active service.
Typed entities and relationships have a long research history, as does the distinction between when a fact is valid and when a database records it.
The program separates entities, relationships, identifiers, and reporting contexts. Its national bitemporal and comparability layers remain program design rather than completed coverage.
Abstention research asks when a model should withhold an answer instead of returning an unsupported one.
These components preserve unknown, held, unavailable, and conflicting states. The page does not report an evaluated model-abstention method or benchmark.
AHRQ publishes a dated health-system universe and separate linkage files with explicit definitions and technical cautions.
The metrics layer preserves that release as a snapshot, including exact identifiers and its documented boundaries. It is not treated as a live national registry.
Current resource boundary · August 8, 2026
What a reader can inspect now.
Access states distinguish resources that can be inspected anonymously from source reviewed through owner access; they describe present availability rather than evidentiary weight.
Acquisition and evidence
Healthcare Data MCP source, documentation, and the public evidence-bundle contract.
An integrated performance evaluation is part of the next research phase.
Institutional identity
FAST Identity source and documentation were reviewed through owner access on August 8, 2026.
This is a completed research record; source access currently requires authentication.
Specialist workups
Healthcare Agents source, routed workups, evidence-pack formats, and review protocols.
Professional judgment and release authority remain with qualified reviewers.
Governed reference
Toolkit source access requires authentication; national coverage is being developed in stages.
Scholarship and distribution
AJHCS retains its own scholarly judgment and release authority.
Temporal and comparability phases
The program design specifies valid and recorded time, reporting contexts, and deterministic comparison policy.
These phases form the next stage of the national evidence layer.
Retrieval needs its own boundary
Healthcare Data MCP grew from the failed profile; MCP stands for Model Context Protocol, the interface used here to give a model access to bounded tools, and the project now provides source-specific acquisition and normalization for public healthcare data. Its tools return structured responses, source metadata, recovery guidance, and local cache behavior, moving repeatable retrieval out of a model's conversational memory and into software that can be inspected and tested.
The public ushso.public-evidence-bundle.v1 contract illustrates the approach: it can carry a source, vintage, producer commit, artifact checksum, coverage state, and conflict state with an observation, making the custody trail available to later identity, metric, profile, and review work instead of leaving it implicit in a conversation. The record supports inspection rather than certifying the observation by itself.
The health-system metrics surface preserves the Agency for Healthcare Research and Quality (AHRQ) 2023 Compendium as a dated snapshot, returning bounded system and hospital-linkage records while keeping that historical frame separate from possible current overlays from the Centers for Medicare & Medicaid Services (CMS). Exact identifiers retain leading zeroes, ambiguous names return candidates rather than a fabricated match, and missing or conflicting observations remain typed states; this follows the source's own caution that spreadsheet handling can alter identifiers without turning the AHRQ release into a live national registry.
Public healthcare data nevertheless remains incomplete: agencies publish on different schedules, names and identifiers drift, and upstream files change shape; a CMS record may describe a certified facility while an audited statement describes a consolidated reporting entity, and neither necessarily equals the health system's public brand. A connector moves the research boundary without erasing it.
Identity, interpretation, and release are different decisions
Retrieval can return the correct row and still support the wrong conclusion: an Employer Identification Number (EIN) identifies a legal filer, a CMS Certification Number (CCN) identifies a Medicare-certified provider, an affiliation does not necessarily establish ownership, and a shared name does not establish that two records describe the same enterprise. When these categories collapse into a single field called “organization,” technically accurate retrieval can become analytically false.
The completed FAST Identity research implementation addresses a deliberately narrow part of that problem, although its source repository is not currently available to anonymous readers; in the reviewed implementation, names alone cannot merge records, aliases may resolve to an existing identity but cannot create or rename one, and lifecycle relationships such as subsystem_of and succeeded_by remain distinct. Form 990 and Schedule R records remain source-attributed observations about legal entities rather than automatic evidence of health-system membership or control.
Public admission was designed to fail closed, requiring separate evidence for whole-system scope, currentness, nonprofit status, in-scope headquarters, and a current subject-bound name; unknown, held, excluded, and review states were legitimate outcomes. FAST Identity is now retained as a completed research record rather than an active identity service, so I use it here as implementation evidence without presenting it as a running national registry.
Identity still does not decide which question matters or who may act on an answer; the public Healthcare Agents project therefore organizes work as routed specialist workups, offline evidence packs, citation-status records, and review protocols, while its evaluator explicitly avoids professional judgments. Human reviewers retain competence checks, dispute resolution, adjudication, and release authority, so these workflows provide vocabulary and evidence expectations without turning a prompt into a credentialed professional.
A reference record is a governed projection
The United States Health Systems Observatory, or USHSO, is the reader-facing reference layer, intended to maintain health-system facts, definitions, evidence, missingness, and interpretation so that a correction can be made in a governed record rather than lost inside an old report; the public site is available, although the underlying Healthcare Toolkit repository is currently private to anonymous readers.
The reviewed Toolkit implementation treats a profile as a projection over typed records: a catalog defines what a metric means, its scope, and its source policy; a metric record stores either a value that fits that definition or an explicit missingness state; source records retain reusable citation metadata; and knowledge objects hold facts, definitions, interpretations, implications, comparison points, and open questions that do not belong in a numeric field. A review gate sits between the workbench and the public profile.
Historical generated-report artifacts for the Jefferson Health profile remain in the audit trail but are retired as the active data-entry contract; the current design instead uses catalog-backed metrics, reusable sources, typed missingness, and supporting knowledge objects, a structure intended to make correction and reuse possible. A public end-to-end trace showing that propagation across every layer is part of the practical roadmap below.
Current program plans extend the model toward valid and recorded time, exact filing-object custody, source-native financial and operating observations, reviewed reporting contexts, and deterministic comparison policy; these are applications of established provenance, temporal-data, and knowledge-representation ideas to the health-system problem, and they define the next stage of national coverage.
Three institutions, distinct responsibilities
AJHCS, USHSO, and Open-Informatics reinforce one another without becoming one chain of command: AJHCS provides a public home for healthcare strategy scholarship, USHSO maintains reference records, definitions, and research tools, and Open-Informatics investigates failures in acquisition, identity, provenance, evaluation, and research software.
Published work exposes unanswered questions, weak definitions, missing sources, and claims that readers cannot yet reproduce; repeated questions can become USHSO records and coverage requirements, allowing Open-Informatics to investigate the technical failures underneath them. A method that becomes stable may then return to USHSO as maintained infrastructure and to AJHCS as a stronger basis for scholarship.
That loop does not transfer authority: Open-Informatics designed and maintains the current AJHCS publishing system, but AJHCS retains authority over its scholarly decisions, while USHSO can provide a governed evidence record without dictating what AJHCS must conclude. Software ownership does not become editorial, source, or professional authority.
A practical research roadmap
The next phase is to connect the program's existing evidence in a public worked profile, beginning with a trace that follows selected observations from source acquisition through identity, interpretation, review, and publication; Healthcare Data MCP and Healthcare Agents already provide public foundations for that work, FAST Identity supplies a completed body of identity research, and USHSO provides the public reference surface through which the integrated method can be examined.
From there, I plan to broaden state and source coverage while adding valid and recorded time, exact filing-object custody, reviewed reporting contexts, and deterministic comparison policy; this work matters because public releases lag institutional change, organizational relationships require time-bounded evidence and adjudication, and legal filers, audited reporting entities, CMS-certified facilities, market brands, and health-system perimeters often fail to coincide. Financial observations may each be accurate and still be incomparable when their subject, period, unit, accounting basis, amendment status, or reporting perimeter differs, which a name match or a similar number cannot settle.
As the architecture develops, language models will remain useful for discovery, candidate generation, extraction assistance, drafting, and critique; lower token prices may make those activities easier to scale, but identity creation, reporting-perimeter decisions, facility-to-system transformations, comparisons, and scholarly release should continue to pass through explicit contracts and accountable human review.
I created Open-Informatics because students and researchers should be able to ask serious questions of complex institutions without first buying an enterprise data contract or becoming a specialist in every source system; the roadmap therefore remains deliberately practical, centered on making more questions inspectable, publishing a worked example of the full path, evaluating it within clear bounds, and preserving the reasons an answer may need further review.
Notes
- Healthcare Data MCP, “Health System Profiler MCP Server: Design Document” (March 2, 2026). This first-party design record documents the Jefferson Health and Lehigh Valley Health Network profiling failure and the proposed source architecture. Its before-and-after table is a design target, not an independent performance evaluation. The file is no longer present on the repository's current branch, so the cited commit preserves the contemporaneous record. Read the preserved design record. ↩
- Public project documentation reviewed August 8, 2026. The Healthcare Data MCP repository documents the source receipts, public evidence contract, recovery behavior, and producer-authority limits described here. The Healthcare Agents repository documents routed workups, offline evidence packs, review protocols, citation-status records, and human-authority boundaries. Volatile repository, release, tool, and coverage counts are intentionally omitted.
- Restricted implementation documentation reviewed August 8, 2026. FAST Data and Healthcare Toolkit were reviewed through the owner's connected GitHub access. Both repositories were private to anonymous readers on the review date. Their documentation grounds the implementation descriptions on this page, but it does not provide public independent verification. No private data or credentials were used.
- Agency for Healthcare Research and Quality. Compendium of U.S. Health Systems, 2023, with system and linkage files plus technical documentation. The page defines its health-system universe, identifies the source files, and warns that common spreadsheet handling can drop leading zeroes from linkage identifiers.
- Provenance and reusable research objects. Wilkinson, M. D. et al., “The FAIR Guiding Principles for scientific data management and stewardship”, Scientific Data 3, 160018 (2016), doi:10.1038/sdata.2016.18. The W3C PROV-O Recommendation supplies the Entity, Activity, and Agent vocabulary used here only as adjacent research context. Open-Informatics does not claim formal conformance.
- Identity and structured knowledge. Christophides, V. et al., “End-to-End Entity Resolution for Big Data: A Survey”, ACM Computing Surveys 53 (2021), also available as arXiv:1905.06397. Ji, S. et al., “A Survey on Knowledge Graphs: Representation, Acquisition, and Applications”, arXiv:2002.00388, later published in IEEE Transactions on Neural Networks and Learning Systems.
- Time and abstention. Snodgrass, R. T. and Ahn, I., “A Taxonomy of Time in Databases”, ACM SIGMOD Record 14, 236-246 (1985). Cai, L. et al., “A Survey on Temporal Knowledge Graph: Representation Learning and Applications”, arXiv:2403.04782. Wen, B. et al., “Know Your Limits: A Survey of Abstention in Large Language Models”, Transactions of the Association for Computational Linguistics 13, 529-556 (2025), doi:10.1162/tacl_a_00754; preprint arXiv:2407.18418. These sources position the design; they do not validate its implementation.