HARPE full DPR library in vitro with Sarkosyl
Identification of the human DPR core promoter element using machine learningA HARPE plasmid library with the 19-nucleotide DPR segment randomized from +17 to +35 relative to the initiator +1 TSS was transcribed in vitro with HeLa nuclear extract. Sarkosyl was added to 0.2% (w/v) 20 seconds after rNTP initiation to restrict transcriptional progression, and RNA/cDNA and plasmid DNA were sequenced for each variant.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
In vitro transcription with HeLa nuclear extract; 0.2% Sarkosyl added 20 seconds after rNTP initiation
Custom HARPE assay based on a modified SuRE episomal plasmid with a randomized DPR in an invariant promoter cassette. Twelve standard reactions were pooled per sample; Sarkosyl was used as a single-round transcription condition. activity_score is RNA RPM divided by DNA RPM, and the extract source is HeLa nuclear extract.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (16 of 16)
| Row | ||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ||||||||||||||||
| 2 | ||||||||||||||||
| 3 | ||||||||||||||||
| 4 | ||||||||||||||||
| 5 | ||||||||||||||||
| 6 | ||||||||||||||||
| 7 | ||||||||||||||||
| 8 | ||||||||||||||||
| 9 | ||||||||||||||||
| 10 | ||||||||||||||||
| 11 | ||||||||||||||||
| 12 | ||||||||||||||||
| 13 | ||||||||||||||||
| 14 | ||||||||||||||||
| 15 | ||||||||||||||||
| 16 | ||||||||||||||||
| 17 | ||||||||||||||||
| 18 | ||||||||||||||||
| 19 | ||||||||||||||||
| 20 | ||||||||||||||||
| 21 | ||||||||||||||||
| 22 | ||||||||||||||||
| 23 | ||||||||||||||||
| 24 | ||||||||||||||||
| 25 | ||||||||||||||||
| 26 | ||||||||||||||||
| 27 | ||||||||||||||||
| 28 | ||||||||||||||||
| 29 | ||||||||||||||||
| 30 | ||||||||||||||||
| 31 | ||||||||||||||||
| 32 | ||||||||||||||||
| 33 | ||||||||||||||||
| 34 | ||||||||||||||||
| 35 | ||||||||||||||||
| 36 | ||||||||||||||||
| 37 | ||||||||||||||||
| 38 | ||||||||||||||||
| 39 | ||||||||||||||||
| 40 | ||||||||||||||||
| 41 | ||||||||||||||||
| 42 | ||||||||||||||||
| 43 | ||||||||||||||||
| 44 | ||||||||||||||||
| 45 | ||||||||||||||||
| 46 | ||||||||||||||||
| 47 | ||||||||||||||||
| 48 | ||||||||||||||||
| 49 | ||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 16 definitions
- element_id
- Package-assigned stable identifier for the tested sequence.
- tested_sequence
- 19-nucleotide randomized DPR sequence corresponding to positions +17 through +35.
- sequence_length
- Length of tested_sequence in nucleotides.
- dna_count_rep1
- Published plasmid DNA read count for biological replicate 1.
- rna_count_rep1
- Published reporter RNA/cDNA read count for biological replicate 1.
- dna_rpm_rep1
- Published plasmid DNA read count normalized to reads per million for replicate 1.
- rna_rpm_rep1
- Published reporter RNA/cDNA read count normalized to reads per million for replicate 1.
- activity_score_rep1
- Transcription strength for replicate 1, calculated as RNA RPM divided by DNA RPM.
- dna_count_rep2
- Published plasmid DNA read count for biological replicate 2.
- rna_count_rep2
- Published reporter RNA/cDNA read count for biological replicate 2.
- dna_rpm_rep2
- Published plasmid DNA read count normalized to reads per million for replicate 2.
- rna_rpm_rep2
- Published reporter RNA/cDNA read count normalized to reads per million for replicate 2.
- activity_score_rep2
- Transcription strength for replicate 2, calculated as RNA RPM divided by DNA RPM.
- n_replicates
- Number of biological replicates represented in the row (2).
- activity_score_mean
- Arithmetic mean of the two replicate activity scores.
- activity_score_sd
- Sample standard deviation of the two replicate activity scores.
Quality control
The GEO inputs had already been filtered by exact flanking-sequence and expected-length read matching, removal of likely index-contamination/invariant reads, and DNA count >=10 with DNA RPM >=0.75. This package re-applied those thresholds, retained finite A/C/G/T sequences, matched the two biological replicates by exact tested_sequence, and retained only sequence elements detected in both replicates; zero-RNA elements were retained as informative inactive measurements.
Curation notes
HARPE is a custom high-throughput randomized promoter assay rather than a conventional barcode-per-oligo MPRA; the sequence itself identifies the downstream variant. Source files are GSM4144996 and GSM4144997 from GEO GSE139635. Scores are relative RNA/DNA ratios, not log2 fold changes or allele-specific effects; the library is synthetic and has no genomic rsIDs or coordinates. HeLa nuclear extract is the biological source; the assay is cell-free.