Experiment / E1CQATQJKPromoter / Core Promoter MPRA

K562 episomal TSS-MPRA, long insert library

Combining TSS-MPRA and sensitive TSS profile dissimilarity scoring to study the sequence determinants of transcription initiation

350-nt construct containing a 303-bp genomic insert library was assayed in K562 cells under basal conditions using promoter / core promoter mpra. The table summarizes element-level reporter activity, barcode-supported RNA/DNA ratios, nucleotide-resolution TSS profiles, and WIP comparison to the paired endogenous csRNA-seq profile.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated K562 cells

Transient episomal TSS-MPRA in K562 cells using a pTSS-MPRA eGFP reporter. Each 350-nt oligonucleotide contains a 303-bp human genomic sequence, an 11-nt transcribed barcode, and cloning/primer arms; the long library was represented by two barcode replicates per element. RNA 5-prime ends report nucleotide-resolution initiation, while barcode DNA counts provide the input normalization.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (47 of 47)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 47 definitions
element_id
Unique tested sequence identifier from the GEO library annotation workbook, with barcode suffix removed.
design_category
Sequence design class inferred from the library identifier: random regulatory element, motif mutation/control, GWAS/SNP allele construct, or splice-donor construct.
variant_id
dbSNP rs identifier parsed from the element identifier or annotation; blank for constructs without an rs identifier.
chromosome
Chromosome reported in the GEO library annotation workbook.
start_grch38
Genomic start coordinate reported in the GEO hg38 annotation workbook; source coordinate convention is preserved.
end_grch38
Genomic end coordinate reported in the GEO hg38 annotation workbook; source coordinate convention is preserved.
strand
Genomic strand reported for the tested sequence.
tss_genomic_position
Main endogenous TSS coordinate parsed from the annotation description.
tss_offset_in_insert_0based
Zero-based position of the annotated main TSS within the genomic insert, parsed from tssOffset.
sequence
Genomic sequence cloned into the reporter, excluding the barcode and cloning arms.
sequence_length
Length in nucleotides of the genomic sequence column.
full_construct_length
Length in nucleotides of the complete synthesized insert construct, including barcode and cloning/primer arms.
annotation_desc
Original compact annotation string containing annotated TSS, motif, mutation, and interval information.
barcode_ids
JSON array of barcode replicate labels retained after DNA-support and barcode-outlier filtering.
n_barcodes_used_rep1
Number of retained barcode replicates with positive DNA abundance in biological replicate 1.
n_barcodes_used_rep2
Number of retained barcode replicates with positive DNA abundance in biological replicate 2.
n_positive_rna_barcodes_rep1
Number of retained barcode replicates with positive RNA abundance in biological replicate 1.
n_positive_rna_barcodes_rep2
Number of retained barcode replicates with positive RNA abundance in biological replicate 2.
n_barcode_outliers_rep1
Number of barcode replicates excluded in biological replicate 1 by the paper-inspired three-global-SD activity outlier rule.
n_barcode_outliers_rep2
Number of barcode replicates excluded in biological replicate 2 by the paper-inspired three-global-SD activity outlier rule.
rna_count_sum_rep1
Sum of released RNA barcode abundances for retained barcodes in biological replicate 1.
dna_count_sum_rep1
Sum of released DNA barcode abundances for retained barcodes in biological replicate 1.
rna_count_sum_rep2
Sum of released RNA barcode abundances for retained barcodes in biological replicate 2.
dna_count_sum_rep2
Sum of released DNA barcode abundances for retained barcodes in biological replicate 2.
rna_cpm_sum_rep1
Retained RNA abundance in counts per million using the complete released RNA block total for biological replicate 1.
dna_cpm_sum_rep1
Retained DNA abundance in counts per million using the complete released DNA block total for biological replicate 1.
rna_cpm_sum_rep2
Retained RNA abundance in counts per million using the complete released RNA block total for biological replicate 2.
dna_cpm_sum_rep2
Retained DNA abundance in counts per million using the complete released DNA block total for biological replicate 2.
activity_log2_rna_dna_rep1
Log2 of CPM-normalized RNA/DNA activity for biological replicate 1.
activity_log2_rna_dna_rep2
Log2 of CPM-normalized RNA/DNA activity for biological replicate 2.
activity_log2_rna_dna_mean
Mean of the two biological-replicate log2 RNA/DNA activity scores.
activity_log2_rna_dna_rep2_minus_rep1
Biological replicate 2 minus replicate 1 log2 RNA/DNA activity score.
activity_rna_dna_ratio_geometric_mean
Geometric mean of the two biological-replicate CPM-normalized RNA/DNA activity ratios.
tss_dominant_position
One-based position within the full synthesized construct with the highest mean MPRA initiation frequency.
tss_dominant_offset_bp
Dominant MPRA initiation position relative to the annotated endogenous TSS; negative values are upstream.
tss_dominant_fraction
Mean MPRA initiation frequency at the dominant position.
tss_profile_entropy_bits
Shannon entropy in bits of the mean normalized MPRA TSS profile; larger values indicate a broader profile.
tss_profile_replicate_pearson
Pearson correlation between the two biological-replicate mean MPRA TSS profiles.
wip_to_csrna_rep1
Wishbone WIP dissimilarity score between MPRA biological replicate 1 and the paired endogenous csRNA-seq profile; lower values indicate greater similarity.
wip_to_csrna_rep2
Wishbone WIP dissimilarity score between MPRA biological replicate 2 and the paired endogenous csRNA-seq profile; lower values indicate greater similarity.
wip_to_csrna_mean
Mean of the two MPRA-to-endogenous csRNA-seq WIP dissimilarity scores.
tss_profile_rep1
JSON array of the biological-replicate-1 mean normalized MPRA TSS frequencies for full-construct positions 1..N.
tss_profile_rep2
JSON array of the biological-replicate-2 mean normalized MPRA TSS frequencies for full-construct positions 1..N.
tss_profile_mean
JSON array of the mean normalized MPRA TSS frequencies for full-construct positions 1..N.
csrna_profile_rep1
JSON array of the paired endogenous csRNA-seq normalized TSS frequencies for positions 1..N, included as the paper's ground-truth reference profile.
csrna_profile_rep2
JSON array of the second paired endogenous csRNA-seq normalized TSS frequencies for positions 1..N.
qc_pass
True for every exported row; only element groups passing the documented barcode and DNA-support QC are present.

Quality control

The authors' released GEO quantification contains per-barcode TPM-transformed, sum-to-one TSS profiles and RNA/DNA barcode abundances; the paper reports high RNA/DNA replicate reproducibility and describes a three-standard-deviation barcode-replicate outlier rule. For this package, barcode rows with non-positive DNA abundance were excluded, barcode activity outliers at or beyond three global residual standard deviations from their element median were excluded, and an element was exported only when at least two retained barcodes had positive DNA in each biological replicate and at least one retained barcode had positive RNA in each replicate. 241 of 250 element groups passed these filters.

Curation notes

The corresponding GEO norm.quant file stores four equal blocks: MPRA RNA/DNA replicate 1, MPRA RNA/DNA replicate 2, endogenous csRNA-seq replicate 1, and endogenous csRNA-seq replicate 2. Only the first two blocks are treated as reporter measurements here; the last two are retained as paired endogenous reference profiles for WIP scoring. Element activity is recomputed from released barcode counts as log2(CPM RNA / CPM DNA), while TSS profiles use the released position-normalized values. The study also tested a small set of 450-, 700-, and 950-bp constructs, but no separate element-level quantification table for those targeted constructs was released in the GEO series; they are therefore not represented as standalone child experiments.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.