Experiment / E3RAVR7KOIntegrated lentiMPRA

K-562 integrated lentiMPRA — Library C with minimal promoter

Design principles of cell-state-specific enhancers in hematopoiesis

The broad Library C pairwise motif library covering 42 transcription factors was delivered to human K-562 cells using the standard minimal promoter and measured in duplicate at the early post-transduction time point.

Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.

Perturbation & assay details

Basal / Untreated

Lentiviral reporter assay using pLS-SceI (Addgene plasmid #137725), with the regulatory element cloned upstream of the indicated minimal promoter and an EGFP reporter. K-562 cells were transduced in duplicate and assayed by DNA/RNA UMI counts from the reporter 5′ UTR. The two replicate activities are log2 library-size-normalized RNA/DNA values; random-DNA subtraction and Trp53 scaling follow the authors' published definitions. The source processing table describes the early K-562 Library C screen as 3 days post-infection with minimum two association reads, UMI read threshold 1, capped DNA UMIs, RNA-only barcodes counted on DNA, minimum 40 UMIs per GRE, and DNA normalization.

Processed data

50 rows per page. Click a cell to inspect its full value.

Visible columns (31 of 31)
Row
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50

Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.

Column dictionary · 31 definitions
element_id
Unique identifier for the tested gene-regulatory element (source CRS).
source_data_object
Name of the main DATA object in the public Figshare R archive.
cell_state_id
Source cell-state cluster identifier; aggregate_HSPC denotes the supplied across-cell-state aggregate view.
cell_state
Cell-state label mapped from the archived cellstate.map vector; K562 denotes State_9K.
library
Paper/library identity assigned from the public DATA-object name and the article (A–H).
source_library_label
Verbatim Library field in the R archive, retained for provenance; it is inconsistent with the object/paper identity for some libraries.
sequence
Reporter insert DNA sequence; in synthetic libraries uppercase letters encode placed motifs and lowercase letters encode background DNA as supplied.
rna_count_rep1
Raw UMI-derived reporter RNA molecule count for replicate 1.
dna_count_rep1
Raw UMI-derived plasmid DNA molecule count for replicate 1.
rna_count_rep2
Raw UMI-derived reporter RNA molecule count for replicate 2.
dna_count_rep2
Raw UMI-derived plasmid DNA molecule count for replicate 2.
rna_normalized_rep1
Library-size-normalized reporter RNA count for replicate 1.
dna_normalized_rep1
Library-size-normalized plasmid DNA count for replicate 1.
rna_normalized_rep2
Library-size-normalized reporter RNA count for replicate 2.
dna_normalized_rep2
Library-size-normalized plasmid DNA count for replicate 2.
activity_log2_raw_rep1
Raw log2 RNA/DNA activity for replicate 1 after library-size normalization.
activity_log2_raw_rep2
Raw log2 RNA/DNA activity for replicate 2 after library-size normalization.
activity_log2_adjusted_rep1
Replicate-1 activity after subtracting the median random-DNA baseline.
activity_log2_adjusted_rep2
Replicate-2 activity after subtracting the median random-DNA baseline.
activity_log2_raw_mean
Mean raw log2 RNA/DNA activity across the two replicates.
activity_log2_adjusted_mean
Mean baseline-adjusted log2 RNA/DNA activity across the two replicates; preferred quantitative activity score.
activity_scaled_mean
Source visualization scale with random-DNA activity at 0 and the Trp53 reference at 1; not the preferred score for modeling or statistical testing.
motif_spacing_bp
Spacing between placed motif sites in base pairs (source spacer).
tf1_name
Name of the first TF motif from the 5′ end of the sequence.
tf1_affinity
Designed affinity of the first TF motif/site set.
tf1_orientation
Orientation of the first TF motif/site set.
tf2_name
Name of the second TF motif from the 5′ end of the sequence.
tf2_affinity
Designed affinity of the second TF motif/site set.
tf2_orientation
Orientation of the second TF motif/site set.
sites_per_factor
Number of sites assigned to each factor in the pair design (source TFnumber).
tf_site_order
Arrangement of paired sites, such as Alternate or Block (source TForder).

Quality control

The authors filtered the GRE–barcode association by alignment score (290–292), dominant assignment support (assigned reads at least five times deviant-assignment reads), barcode homopolymer content (removed barcodes with more than 10 identical nucleotides), and sequencing-error correction. MPRA reads required concordant forward/reverse barcode reads and underwent UMI error correction. The supplied DATA frame is the authors' post-processing/post-QC element-by-cell-state table. Package QC additionally required complete identifiers, non-empty sequences, positive raw DNA/RNA counts, and finite normalized/activity values. The public DATA object retained 23,392 of 23,392 rows (0 removed); the adjusted-activity replicate Pearson correlation is 0.977.

Curation notes

The DATA object is K562.libC.minP.tra and its source_library_label is LibB. The object name and pair-design schema are used for the paper library identity.

Cite OpenMPRA

Cite the OpenMPRA database. Include your access date because the collection changes over time.

Please also cite the source studies when using their data.