Experiment / E3MHJA30LIntegrated lentiMPRA

HepG2 ENCODE pilot lentiMPRA of non-B DNA-associated regulatory elements

High-throughput characterization of the role of non-B DNA motifs on promoter function

Three isogenic HepG2 replicates tested a library of 200-bp genomic elements in an integrated lentiMPRA reporter. RNA/DNA activity measurements are joined to the original ENCODE reference sequences and annotated for non-B DNA motif features.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

ENCODE functional characterization experiment ENCSR463IRX; three isogenic HepG2 replicates using an upstream lentiMPRA reporter with the test element and a barcode in the 5' UTR. The processed source quantification is ENCODE file ENCFF285OWV and contains DNA count, RNA count, RNA/DNA ratio, log2 activity, and observed-barcode fields.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (39 of 39)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 39 definitions
element_id
Element name from the ENCODE combined quantification table.
chromosome
GRCh38 chromosome from the ENCODE activity BED; blank when no coordinate was provided.
start_grch38_0based
0-based inclusive GRCh38 start coordinate from the ENCODE activity BED; blank when unavailable.
end_grch38_0based_exclusive
0-based exclusive GRCh38 end coordinate from the ENCODE activity BED; blank when unavailable.
strand
Strand from the ENCODE activity BED; blank when unavailable.
sequence_200bp
The original 200-bp ENCODE reference sequence in reporter-design orientation.
sequence_length_bp
Length of sequence_200bp in base pairs.
rep1_dna_count
ENCODE normalized plasmid/library DNA abundance for replicate 1.
rep1_rna_count
ENCODE normalized reporter RNA abundance for replicate 1.
rep1_activity_ratio
ENCODE RNA/DNA activity ratio for replicate 1.
rep1_log2_activity
ENCODE log2 RNA/DNA activity score for replicate 1.
rep1_n_observed_barcodes
Number of observed barcodes contributing to replicate 1 quantification.
rep2_dna_count
ENCODE normalized plasmid/library DNA abundance for replicate 2.
rep2_rna_count
ENCODE normalized reporter RNA abundance for replicate 2.
rep2_activity_ratio
ENCODE RNA/DNA activity ratio for replicate 2.
rep2_log2_activity
ENCODE log2 RNA/DNA activity score for replicate 2.
rep2_n_observed_barcodes
Number of observed barcodes contributing to replicate 2 quantification.
rep3_dna_count
ENCODE normalized plasmid/library DNA abundance for replicate 3.
rep3_rna_count
ENCODE normalized reporter RNA abundance for replicate 3.
rep3_activity_ratio
ENCODE RNA/DNA activity ratio for replicate 3.
rep3_log2_activity
ENCODE log2 RNA/DNA activity score for replicate 3.
rep3_n_observed_barcodes
Number of observed barcodes contributing to replicate 3 quantification.
mean_dna_count
Arithmetic mean of the three replicate DNA counts.
mean_rna_count
Arithmetic mean of the three replicate RNA counts.
mean_activity_ratio
Arithmetic mean of the three replicate RNA/DNA activity ratios.
mean_log2_activity
Arithmetic mean of the three replicate log2 activity scores.
sd_log2_activity
Sample standard deviation of the three replicate log2 activity scores.
min_n_observed_barcodes
Minimum observed-barcode count across the three replicates.
mean_n_observed_barcodes
Arithmetic mean of observed-barcode counts across the three replicates.
g4_forward_count
Count of published consensus G-quadruplex regex matches on the forward sequence.
g4_reverse_complement_count
Count of published consensus G-quadruplex regex matches on the reverse-complement sequence.
g4_count
Total G-quadruplex regex matches across forward and reverse-complement sequences.
z_dna_count
Count of qualifying alternating purine/pyrimidine tracts of at least 10 bp, excluding AT and TA pairs.
inverted_repeat_count
Count of exact 10-bp-arm inverted-repeat matches with 0-4 bp spacers.
direct_repeat_count
Count of exact 10-bp-arm direct-repeat matches with 0-4 bp spacers, excluding highly tandem/repetitive arms.
mirror_repeat_count
Count of exact 10-bp-arm mirror-repeat matches with 0-4 bp spacers.
short_tandem_repeat_count
Count of short tandem-repeat runs with 1-9 bp repeat units, at least 5 copies, and at least 12 bp total length.
h_dna_count
Count of high-AG (>90%) 10-bp-arm mirror-repeat matches with 0-7 bp spacers, a sequence-level H-DNA proxy based on the paper definition.
qc_pass
TRUE for every row retained after the package-level QC filters.

Quality control

The study/MPRAflow workflow used consensus insert-barcode processing, strict sequence matching, UMI handling, and removal of elements with fewer than 3 associated barcodes. For this package, no_BC was excluded as a technical control; elements were retained only when all three replicates were present, every replicate had n_obs_bc >= 3, all activity/count fields were finite and valid, and a matching 200-bp ACGT-only reference sequence was available. This retained 9,303 of 9,317 non-no_BC elements (13 low-barcode/invalid elements and 1 incomplete element excluded); 9,199 retained elements had a public GRCh38 BED coordinate.

Curation notes

This is the Figure 3 HepG2 ENCODE-linked pilot experiment, not the separate custom HEK-293T/K562 or NPC BioProject library. Motif counts were derived from the deposited 200-bp sequences using the repository's published G4, inverted/direct/mirror-repeat, STR, and Z-DNA rules; H-DNA uses the paper's high-AG mirror-repeat definition with the repository's 10-bp arm convention. Coordinates are supplementary annotations and are blank for elements absent from the current ENCODE GRCh38 BED. The ENCODE BED contains 101 joined intervals that do not span exactly 200 bp (mostly 171 bp), so coordinates should not be used to reconstruct the reporter sequence.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.