Experiment / E82J7WBXUIntegrated lentiMPRA

Chromosomally integrated lentiMPRA of HepG2 candidate liver enhancers (wild-type integrase / WT)

A systematic comparison reveals substantial differences in chromosomal versus episomal encoding of enhancer activity

The 2,440-element 5′ enhancer/3′ barcode library was packaged with wild-type lentiviral integrase and introduced into HepG2 cells for chromosomal integration. Three independent infections were profiled by matched DNA and mRNA barcode sequencing for 2,236 candidate liver enhancers and 204 controls.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / untreated; wild-type-integrase lentiviral delivery

The library used a 5′ 171-bp candidate sequence upstream of a minimal promoter and EGFP, with a 15-bp barcode in the EGFP 3′ UTR; each insert had 100 designed barcodes and the cassette was flanked by antirepressor/SAR elements. HepG2 DNA and mRNA were collected at day 4 after an estimated 50 viral particles per cell; three independent infections were sequenced on Illumina NextSeq.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (53 of 53)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 53 definitions
element_id
Published insert/region identifier from GEO ActivityRatios.
design_region_id
Identifier in the oligo design; differs from element_id only for four controls whose activity labels add a bracketed coordinate span.
category
Broad published design category: type1–type4, SRE, or control.
category_detail
Published subtype or control label, such as HNF4A-ChMod, positive, or negative.
element_sequence
The intended 171-bp enhancer or control sequence extracted from the designed 230-bp oligo.
sequence_length_bp
Length of element_sequence in base pairs.
coordinate_assembly
Assembly label parsed from the published identifier; candidate regions use GRCh37 and controls preserve hg18, hg19, or mm9 labels.
coordinate
High-level target interval parsed from the published region identifier.
chromosome
Chromosome parsed from coordinate.
start
1-based target interval start parsed from coordinate.
end
Target interval end parsed from coordinate.
barcode_count_designed
Number of designed 15-bp reporter barcodes assigned to the insert; the library design used 100.
activity_rep1_rna_dna
Published wild-type-integrase/integrated RNA/DNA activity ratio for replicate 1.
activity_rep2_rna_dna
Published wild-type-integrase/integrated RNA/DNA activity ratio for replicate 2.
activity_rep3_rna_dna
Published wild-type-integrase/integrated RNA/DNA activity ratio for replicate 3.
activity_combined_rna_dna
Published WT activity ratio combined across replicates after replicate-median normalization.
activity_log2_rna_dna
log2 of activity_combined_rna_dna, derived in this package.
replicate_activity_mean_rna_dna
Arithmetic mean of the three published WT replicate ratios; descriptive and not the authors’ combined value.
replicate_activity_sd_rna_dna
Sample standard deviation of the three published WT replicate ratios.
barcodes_common_rep1
Designed barcode IDs observed in both DNA and RNA count files for WT replicate 1.
barcodes_common_rep2
Designed barcode IDs observed in both DNA and RNA count files for WT replicate 2.
barcodes_common_rep3
Designed barcode IDs observed in both DNA and RNA count files for WT replicate 3.
barcodes_common_mean
Arithmetic mean of the three shared RNA/DNA barcode counts.
dna_counts_common_rep1
Sum of assigned DNA counts over barcodes shared by RNA and DNA in WT replicate 1.
dna_counts_common_rep2
Sum of assigned DNA counts over barcodes shared by RNA and DNA in WT replicate 2.
dna_counts_common_rep3
Sum of assigned DNA counts over barcodes shared by RNA and DNA in WT replicate 3.
rna_counts_common_rep1
Sum of assigned RNA counts over barcodes shared by RNA and DNA in WT replicate 1.
rna_counts_common_rep2
Sum of assigned RNA counts over barcodes shared by RNA and DNA in WT replicate 2.
rna_counts_common_rep3
Sum of assigned RNA counts over barcodes shared by RNA and DNA in WT replicate 3.
paired_other_context_activity_combined_rna_dna
Published combined MT/nonintegrating RNA/DNA activity ratio for the paired context.
integrated_minus_episomal_log2_ratio
log2(WT combined activity / MT combined activity); positive values favor the integrated context.
qc_pass
TRUE for rows retained after package-level design, activity, and shared-barcode QC.
OC
Majority HepG2 OpenChromatin state annotation.
SegWay
Majority HepG2 SegWay genome-segmentation annotation.
ChromHMM7
Majority HepG2 7-state ChromHMM annotation.
ChromHMM15
Majority Roadmap HepG2 15-state ChromHMM annotation.
ERBFeature
Majority Ensembl Regulatory Build regulatory-feature annotation.
DHS
Number of element bases overlapping HepG2 DNase hypersensitivity calls.
H3K27ac
Number of element bases covered by HepG2 H3K27ac ChIP-seq peaks.
H3K27acAve
Average HepG2 H3K27ac peak signal over the element.
H3K27acMax
Maximum HepG2 H3K27ac peak signal over the element.
TFcount
Number of distinct assayed transcription factors with ChIP-seq peak overlap.
EP300
Number of element bases covered by HepG2 EP300 ChIP-seq peaks.
FOXA1
Number of element bases covered by HepG2 FOXA1 ChIP-seq peaks.
FOXA2
Number of element bases covered by HepG2 FOXA2 ChIP-seq peaks.
HNF4A
Number of element bases covered by HepG2 HNF4A ChIP-seq peaks.
RAD21
Number of element bases covered by HepG2 RAD21 ChIP-seq peaks.
CHD2
Number of element bases covered by HepG2 CHD2 ChIP-seq peaks.
SMC3
Number of element bases covered by HepG2 SMC3 ChIP-seq peaks.
GC
CADD v1.3 average local GC percentage annotation.
CpG
CADD v1.3 average local CpG percentage annotation.
minDistTSS
CADD v1.3 distance to the closest transcribed-sequence start site.
CADD
CADD v1.3 PHRED-scaled sequence annotation.

Quality control

Author QC was retained: 15-bp barcode consensus sequences were matched to the designed barcode set without clustering; only barcodes observed in both RNA and DNA in the same sample were used; counts were normalized to counts per million; per-insert RNA/DNA ratios were calculated from summed barcode counts; replicate ratios were divided by their replicate median before averaging. Package-level integrity QC additionally required an exact design match, 100 designed barcodes, positive numeric activity in all three replicates, and at least 3 shared RNA/DNA barcodes per replicate. All 2,440/2,440 inserts passed; the minimum shared-barcode count was 9, 9, and 8 across the three WT replicates.

Curation notes

This is the wild-type-integrase arm of a paired WT/MT lentiMPRA and represents random chromosomal integration of the reporter library. The study is region-focused rather than variant-focused: it tests annotated 171-bp candidate liver enhancers and controls, with no designed allele contrasts. ActivityRatios contains all 2,440 inserts; source values marked NA in selected annotation fields are emitted as blank cells, while the primary activity fields are complete. The four endogenous/SRE controls include hg19, hg18, or mm9 labels in their identifiers and are retained as controls.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.