Live-cell imaging of enhancer-promoter dynamics reveals transient contact-driven gene activation
收藏资源简介:
Overview This repository contains all the raw and processed trajectory data associated with “Live-cell imaging of enhancer-promoter dynamics reveals transient contact-driven gene activation”. In this ReadMe file we provide the following information: The cell lines and conditions used in this study A summary of how the data was collected A description of the organization of the live-cell trajectory data Structure of data The underlying mathematical values are identical between the .csv and .npz files. Please choose the format that best fits your analysis pipeline: Option A: Tabular Format (.csv) Best for dataframes (Pandas, R). The data is provided in a flattened format. To maintain direct 1:1 structural parity with the matrix dimensions as Option B below, the .csv files retain the missing frames that should be interpreted as NaNs (shown as empty values in .csv). This allows you to easily reshape the CSV columns back into N x T blocks. Option B: Matrix Format (.npz) Best for tensor operations. By loading the .npz file (e.g., data = np.load(filepath)), you can access the variables via dictionary keys. The arrays are shaped as (N, T, ...) where N is the number of distinct tracks and T is the maximum number of frames. Note on Matrix Padding: Trajectories of unequal lengths or those containing missing frames are padded with NaN values to maintain a uniform N x T rectangular structure. Note on Indexing: The first column of the matrices (index 0) aligns to the start of the trajectory (e.g., the first detected localization). The true absolute acquisition time and frame number for each matrix entry are preserved in the respective time and frame arrays. Cell lines and conditions The dataset contains live-cell imaging trajectories from cell lines with varying E-P genomic distances, CTCF configurations, and degron-targeted depletion conditions, as well as trajectories from control cell lines. To easily match the biological conditions described in the paper with the file names in the dataset, please refer to the mapping table below. (Note: The final .csv/.npz file names are prefixed with the temporal resolution, e.g., 5s_340kb_Ce_Cp_None.csv or 30s_85kb_IAA.npz). Biological Condition Description CSV/NPZ File Base Name 339CECP 339 kb E-P pair with convergent CTCF binding sites 340kb_Ce_Cp_None 339CECP (ΔCTCF) 339CECP treated with dTAG-13 to deplete CTCF 340kb_Ce_Cp_dTAG 339CECP (ΔRAD21) 339CECP treated with 5ph-IAA to deplete RAD21 340kb_Ce_Cp_IAA 339CECP (ΔRAD21 & ΔCTCF) 339CECP treated with both 5ph-IAA and dTAG-13 340kb_Ce_Cp_IAAdTAG 339CE 339 kb E-P pair with CTCF binding sites on enhancer side only 340kb_Ce_None 339CP 339 kb E-P pair with CTCF binding sites on promoter side only 340kb_Cp_None 339noEP 339 kb convergent CTCF binding sites without the E-P pair 340kb_Ce_Cp_noE_noP_None 339noE 339 kb convergent CTCF binding sites with the promoter alone 340kb_Ce_Cp_noE_None 339noC 339 kb E-P pair without CTCF binding sites 340kb_None 339noC (ΔRAD21) 339noC treated with 5ph-IAA to deplete RAD21 340kb_IAA 253noC 253 kb E-P pair without CTCF binding sites 255kb_None 253noC (ΔRAD21) 253noC treated with 5ph-IAA to deplete RAD21 255kb_IAA 170noC 170 kb E-P pair without CTCF binding sites 170kb_None 170noC (ΔRAD21) 170noC treated with 5ph-IAA to deplete RAD21 170kb_IAA 87noC 87 kb E-P pair without CTCF binding sites 85kb_None 87noC (ΔRAD21) 87noC treated with 5ph-IAA to deplete RAD21 85kb_IAA 1.5noC 1.5 kb E-P pair without CTCF binding sites 1.5kb_None 1.5noC (ΔRAD21) 1.5noC treated with 5ph-IAA to deplete RAD21 1.5kb_IAA Sub-proximal label control Control cell line with proximal labels (center-to-center distance of 3.5 kb, 2.4 kb between the labels) 2.362kb_noE_noP_None Proximal label control Control cell line with sub-proximal labels (center-to-center distance of 1.5kb, 0.4 kb between the labels) 0.395kb_noE_noP_None Note: For depletion conditions, cell lines were treated with 100 µM 5ph-IAA (RAD21) or 500 nM dTAG-13 (CTCF) or both for 2 hours prior to the start of imaging. Data and data processing Trajectories were obtained from 3D and 3-color time-series live-cell imaging of mouse embryonic stem cells (mESCs). The synthetic enhancer-promoter pairs were integrated into a clean region on chromosome 2 (mm39: ~178,200,000 to 179,600,000). Acquisition was performed using a ZEISS Lattice Lightsheet 7 (LLS7) microscope. Imaging was performed in three colors simultaneously to track the enhancer (synBsr1-ParB-2x-mScarlet3, 561 nm), the promoter (OR3-2x-mStayGold, 488 nm), and nascent transcription (MCP-2x-HaloTag, 640 nm). Data were collected at three different temporal resolutions: 30-second frame rate: 6 hours total duration (30 ms exposure, with auto-focus). 5-second frame rate: 1 hour total duration (30 ms exposure, without auto-focus). 0.5-second frame rate: 6 minutes total duration (10 ms exposure, without auto-focus). The raw 3D image time series were processed using our custom Fyrtarn framework (available at https://github.com/ahansenlab/Fyrtarn). This pipeline performed nuclear segmentation, FracShift subpixel localization, chromatic aberration correction, and photobleaching correction. The trajectory data provided in this repository has undergone manual QC (to remove replicated and false positive dots) and statistical ensemble filtering (to remove single-frame jumps and extreme localization anomalies). Because of this filtering, some trajectories do not span the entire time range and some have gaps. To ensure broad usability, each dataset is identically exported in two formats: long-format tabular .csv files and compressed NumPy .npz matrices. File names are formatted as: {framerate}_{npz_base_name}.npz or {framerate}_{npz_base_name}.csv (e.g., 5s_170kb_None.npz or 5s_170kb_None.csv). Data Variables 3D Distance: The absolute Euclidean distance between the fluorescent labels of the enhancer and promoter ((Δx² + Δy² + Δz²)0.5) in nanometers. In .csv: Column 3D_distance (nm) In .npz: Key 3D_distance, shaped (N, T) Enhancer coordinates: the absolute x, y, and z coordinates of the enhancer label in nanometers. In .csv: Columns enh_x (nm), enh_y (nm), enh_z (nm) In .npz: Key enhancer_coordinate, shaped (N, T, 3) Promoter Coordinates: the absolute x, y, and z coordinates of the promoter label in nanometers. In .csv: Columns pro_x (nm), pro_y (nm), pro_z (nm) In .npz: Key promoter_coordinate, shaped (N, T, 3) MS2 Intensity: the raw MS2 (nascent transcription) intensity in arbitrary units (au). In .csv: Column intensity (au) In .npz: Key intensity, shaped (N, T) Corrected Intensity: MS2 intensity values that have been corrected for the observed decay post background subtraction and photobleaching correction (see Supplementary Note 3 of the paper for details). In .csv: Column corrected_intensity (au) In .npz: Key corrected_intensity, shaped (N, T) Absolute Time: the absolute imaging time in seconds for the localization. In .csv: Column time (s) In .npz: Key time, shaped (N, T) Absolute Frame: the absolute frame index for the localization. In .csv: Column frame In .npz: Key frame, shaped (N, T) Identifiers: Strings containing the relative file path of the source tracking file, allowing direct trace-back to the raw localization data. In .csv: Column identifier In .npz: Key identifiers, shaped (N)



