K562 ENCODE pilot lentiMPRA of non-B DNA-associated regulatory elements
High-throughput characterization of the role of non-B DNA motifs on promoter functionThree isogenic K562 replicates tested a library of 200-bp genomic elements in an integrated lentiMPRA reporter. RNA/DNA activity measurements are joined to the original ENCODE reference sequences and annotated for non-B DNA motif features.
Processed tables are specific to each experiment. Column names, units, measurements, and table structure are not standardized across the database. Check this experiment’s column definitions and quality-control notes before comparing or combining data.
Perturbation & assay details
Basal / Untreated
ENCODE functional characterization experiment ENCSR460LZI; three isogenic K562 replicates using an upstream lentiMPRA reporter with the test element and a barcode in the 5' UTR. The processed source quantification is ENCODE file ENCFF788GLA and contains DNA count, RNA count, RNA/DNA ratio, log2 activity, and observed-barcode fields.
Processed data
50 rows per page. Click a cell to inspect its full value.
Visible columns (39 of 39)
| Row | |||||||||||||||||||||||||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | |||||||||||||||||||||||||||||||||||||||
| 2 | |||||||||||||||||||||||||||||||||||||||
| 3 | |||||||||||||||||||||||||||||||||||||||
| 4 | |||||||||||||||||||||||||||||||||||||||
| 5 | |||||||||||||||||||||||||||||||||||||||
| 6 | |||||||||||||||||||||||||||||||||||||||
| 7 | |||||||||||||||||||||||||||||||||||||||
| 8 | |||||||||||||||||||||||||||||||||||||||
| 9 | |||||||||||||||||||||||||||||||||||||||
| 10 | |||||||||||||||||||||||||||||||||||||||
| 11 | |||||||||||||||||||||||||||||||||||||||
| 12 | |||||||||||||||||||||||||||||||||||||||
| 13 | |||||||||||||||||||||||||||||||||||||||
| 14 | |||||||||||||||||||||||||||||||||||||||
| 15 | |||||||||||||||||||||||||||||||||||||||
| 16 | |||||||||||||||||||||||||||||||||||||||
| 17 | |||||||||||||||||||||||||||||||||||||||
| 18 | |||||||||||||||||||||||||||||||||||||||
| 19 | |||||||||||||||||||||||||||||||||||||||
| 20 | |||||||||||||||||||||||||||||||||||||||
| 21 | |||||||||||||||||||||||||||||||||||||||
| 22 | |||||||||||||||||||||||||||||||||||||||
| 23 | |||||||||||||||||||||||||||||||||||||||
| 24 | |||||||||||||||||||||||||||||||||||||||
| 25 | |||||||||||||||||||||||||||||||||||||||
| 26 | |||||||||||||||||||||||||||||||||||||||
| 27 | |||||||||||||||||||||||||||||||||||||||
| 28 | |||||||||||||||||||||||||||||||||||||||
| 29 | |||||||||||||||||||||||||||||||||||||||
| 30 | |||||||||||||||||||||||||||||||||||||||
| 31 | |||||||||||||||||||||||||||||||||||||||
| 32 | |||||||||||||||||||||||||||||||||||||||
| 33 | |||||||||||||||||||||||||||||||||||||||
| 34 | |||||||||||||||||||||||||||||||||||||||
| 35 | |||||||||||||||||||||||||||||||||||||||
| 36 | |||||||||||||||||||||||||||||||||||||||
| 37 | |||||||||||||||||||||||||||||||||||||||
| 38 | |||||||||||||||||||||||||||||||||||||||
| 39 | |||||||||||||||||||||||||||||||||||||||
| 40 | |||||||||||||||||||||||||||||||||||||||
| 41 | |||||||||||||||||||||||||||||||||||||||
| 42 | |||||||||||||||||||||||||||||||||||||||
| 43 | |||||||||||||||||||||||||||||||||||||||
| 44 | |||||||||||||||||||||||||||||||||||||||
| 45 | |||||||||||||||||||||||||||||||||||||||
| 46 | |||||||||||||||||||||||||||||||||||||||
| 47 | |||||||||||||||||||||||||||||||||||||||
| 48 | |||||||||||||||||||||||||||||||||||||||
| 49 | |||||||||||||||||||||||||||||||||||||||
| 50 |
Filters apply to this table only. The CSV download contains the complete processed table; filtered rows are available through the API.
Column dictionary · 39 definitions
- element_id
- Element name from the ENCODE combined quantification table.
- chromosome
- GRCh38 chromosome from the ENCODE activity BED; blank when no coordinate was provided.
- start_grch38_0based
- 0-based inclusive GRCh38 start coordinate from the ENCODE activity BED; blank when unavailable.
- end_grch38_0based_exclusive
- 0-based exclusive GRCh38 end coordinate from the ENCODE activity BED; blank when unavailable.
- strand
- Strand from the ENCODE activity BED; blank when unavailable.
- sequence_200bp
- The original 200-bp ENCODE reference sequence in reporter-design orientation.
- sequence_length_bp
- Length of sequence_200bp in base pairs.
- rep1_dna_count
- ENCODE normalized plasmid/library DNA abundance for replicate 1.
- rep1_rna_count
- ENCODE normalized reporter RNA abundance for replicate 1.
- rep1_activity_ratio
- ENCODE RNA/DNA activity ratio for replicate 1.
- rep1_log2_activity
- ENCODE log2 RNA/DNA activity score for replicate 1.
- rep1_n_observed_barcodes
- Number of observed barcodes contributing to replicate 1 quantification.
- rep2_dna_count
- ENCODE normalized plasmid/library DNA abundance for replicate 2.
- rep2_rna_count
- ENCODE normalized reporter RNA abundance for replicate 2.
- rep2_activity_ratio
- ENCODE RNA/DNA activity ratio for replicate 2.
- rep2_log2_activity
- ENCODE log2 RNA/DNA activity score for replicate 2.
- rep2_n_observed_barcodes
- Number of observed barcodes contributing to replicate 2 quantification.
- rep3_dna_count
- ENCODE normalized plasmid/library DNA abundance for replicate 3.
- rep3_rna_count
- ENCODE normalized reporter RNA abundance for replicate 3.
- rep3_activity_ratio
- ENCODE RNA/DNA activity ratio for replicate 3.
- rep3_log2_activity
- ENCODE log2 RNA/DNA activity score for replicate 3.
- rep3_n_observed_barcodes
- Number of observed barcodes contributing to replicate 3 quantification.
- mean_dna_count
- Arithmetic mean of the three replicate DNA counts.
- mean_rna_count
- Arithmetic mean of the three replicate RNA counts.
- mean_activity_ratio
- Arithmetic mean of the three replicate RNA/DNA activity ratios.
- mean_log2_activity
- Arithmetic mean of the three replicate log2 activity scores.
- sd_log2_activity
- Sample standard deviation of the three replicate log2 activity scores.
- min_n_observed_barcodes
- Minimum observed-barcode count across the three replicates.
- mean_n_observed_barcodes
- Arithmetic mean of observed-barcode counts across the three replicates.
- g4_forward_count
- Count of published consensus G-quadruplex regex matches on the forward sequence.
- g4_reverse_complement_count
- Count of published consensus G-quadruplex regex matches on the reverse-complement sequence.
- g4_count
- Total G-quadruplex regex matches across forward and reverse-complement sequences.
- z_dna_count
- Count of qualifying alternating purine/pyrimidine tracts of at least 10 bp, excluding AT and TA pairs.
- inverted_repeat_count
- Count of exact 10-bp-arm inverted-repeat matches with 0-4 bp spacers.
- direct_repeat_count
- Count of exact 10-bp-arm direct-repeat matches with 0-4 bp spacers, excluding highly tandem/repetitive arms.
- mirror_repeat_count
- Count of exact 10-bp-arm mirror-repeat matches with 0-4 bp spacers.
- short_tandem_repeat_count
- Count of short tandem-repeat runs with 1-9 bp repeat units, at least 5 copies, and at least 12 bp total length.
- h_dna_count
- Count of high-AG (>90%) 10-bp-arm mirror-repeat matches with 0-7 bp spacers, a sequence-level H-DNA proxy based on the paper definition.
- qc_pass
- TRUE for every row retained after the package-level QC filters.
Quality control
The study/MPRAflow workflow used consensus insert-barcode processing, strict sequence matching, UMI handling, and removal of elements with fewer than 3 associated barcodes. For this package, no_BC was excluded as a technical control; elements were retained only when all three replicates were present, every replicate had n_obs_bc >= 3, all activity/count fields were finite and valid, and a matching 200-bp ACGT-only reference sequence was available. This retained 7,328 of 7,367 non-no_BC elements (12 low-barcode/invalid elements and 27 incomplete elements excluded); 7,160 retained elements had a public GRCh38 BED coordinate.
Curation notes
This is the Figure 3 K562 ENCODE-linked pilot experiment, not the separate custom HEK-293T/K562 or NPC BioProject library. Motif counts were derived from the deposited 200-bp sequences using the repository's published G4, inverted/direct/mirror-repeat, STR, and Z-DNA rules; H-DNA uses the paper's high-AG mirror-repeat definition with the repository's 10-bp arm convention. Coordinates are supplementary annotations and are blank for elements absent from the current ENCODE GRCh38 BED.