About OpenMPRA
OpenMPRA is a continuously updated database of massively parallel reporter assay (MPRA) studies. It brings publication metadata, experimental conditions, processed tables, and raw source files into one searchable resource.
How studies are added
A system flags papers as potential candidates for MPRA studies. An AI agent then works through the source material and relevant data to package each qualifying study for the database.
A study package contains publication metadata, relevant raw inputs, and a separate processed table for each experiment. The agent records experimental conditions, column definitions, quality-control criteria, and notes about ambiguities or limitations. Information that cannot be established from the source material is left unreported rather than guessed.
The website indexes finalized study packages and periodically checks for additions and revisions. This indexing step makes the packages searchable; it is not an independent scientific review. AI-assisted curation can introduce errors, so researchers should check the source publication and data when interpreting results.
What qualifies as an MPRA study?
The database criteria require a study to generate a pooled or multiplexed reporter-library experiment on many identifiable DNA or RNA sequence units. Reporter output must be assigned to individual units through sequencing or an equivalent high-throughput molecular counting method. The measured outcome must concern a cis-regulatory effect, such as activation, repression, promoter activity, RNA stability, translation, splicing, or silencing.
Approximately 100 or more distinct sequence units is strong practical evidence, rather than a strict threshold. Papers that only reuse existing data, conventional small reporter panels, and endogenous RNA-seq or ATAC-seq without sequence-linked reporter measurements do not qualify on their own.
Understanding the data
Processed tables are experiment-specific
Each experiment has one processed CSV table. Table structure, column names, units, and measurements vary between experiments and are not standardized across the database. Column definitions, QC descriptions, reference assemblies, and curation notes should be consulted before comparing or combining results.
Raw files retain the source inputs
Study pages provide retained supplemental and deposited inputs, available individually or together as a ZIP. These are distinct from the processed experiment tables. Raw sequencing reads are normally omitted when they are not needed for curation. Follow the publication links to obtain the paper’s full text.
Missing values are meaningful
Unknown metadata is represented as null in the API. A missing reference genome may reflect a synthetic library; an unreported identifier is not evidence that a biological feature was absent.
Updates and access
The app refreshes its index in the background, by default every 15 minutes. A successful refresh incorporates new finalized studies, metadata edits, and removals. This is the frequency of indexing, not a claim about when the underlying research data were last changed. If the file server is temporarily unavailable, the last successful index remains available.
The read-only API provides programmatic access to study and experiment metadata, table rows, processed CSVs, and raw files. Read the API documentation.
Citation, reuse, and corrections
Use the Citation link in the header for the database citation, and cite the source studies when using their data. Original source licenses and reuse terms apply; inclusion in OpenMPRA does not assign a new data license.
Report problems or suggest improvements through the GitHub repository. Include the study or experiment ID so the affected record can be identified.