antiCD19_CART_NTX_SCENIC
收藏资源简介:
Single-cell 5′ multiomic profiling of human bone marrow and blood lymphocytes before and after anti CD19 CAR T therapy therapy for HLA desensitizationWe performed five single-cell captures (NTX1–NTX5) on primary human bone marrow and blood using the 10x Genomics Chromium Single Cell 5′ workflow. Library types included 5′ gene expression (GEX), cell-surface protein profiling via CITE-seq (TotalSeq C), and paired V(D)J repertoires (TCR and BCR). Cell inputs per capture were: NTX1 (20,000 plasma cells; BM before therapy), NTX2 (14,000 memory B cells + 4,000 T cells + 2,000 NK; BM before therapy), NTX3 (14,000 plasma cells & memory B cells + 4,000 T + 2,000 NK; blood before therapy), NTX4 (14,000 plasma cells + 4,000 T + 2,000 NK; BM after therapy), and NTX5 (14,000 T + 6,000 NK; blood after therapy). Libraries were quantified and sequenced on an Illumina NextSeq 500 (GEX: High Output v2 150 cycles; VDJ: Mid Output v2 300 cycles). Processed outputs include CellBender-filtered matrices and integrated AnnData (h5ad) objects for joint analyses.Five 10x Genomics Chromium Single Cell 5′ captures (NTX1–NTX5) with matched GEX, CITE-seq (TotalSeq C), and V(D)J libraries (TCR and BCR as applicable). BM vs blood; before vs after therapy design: NTX1 (BM pre), NTX2 (BM pre), NTX3 (blood pre), NTX4 (BM post), NTX5 (blood post). CITE-seq prepared with Feature Barcode Library Kit and TotalSeq C panel; V(D)J enrichment using Human T-cell and B-cell kits; sequenced on Illumina NextSeq 500 brief overview of library generation#Samples represent before-therapy and after-therapy time points; no ex vivo stimulation prior to capture.Cell isolation by density gradient; FACS enrichment of targeted lymphocyte subsets; immediate encapsulation using 10x Genomics Chromium Controller (Single Cell 5′).10x Genomics Single Cell 5′ Library & Gel Bead Kit and Feature Barcode Library Kit. CITE-seq libraries prepared with TotalSeq C antibodies and Single Index Kit N Set A. V(D)J target enrichment using Chromium Single Cell V(D)J Enrichment Kits (Human T cells and B cells). Final libraries quantified by Qubit HS DNA; size distribution assessed with Fragment Analyzer (Agilent) HS NGS Fragment Kit (1–6000 bp). Sequencing performed on Illumina NextSeq 500 (GEX: High Output v2, 150 cycles; VDJ: Mid Output v2, 300 cycles).Demultiplexing with cellranger mkfastq.Gene expression quantified with cellranger count (5′ + Feature Barcoding).V(D)J reconstructed with cellranger vdj for TCR/BCR libraries.Background removal on GEX with CellBender; integration and analysis using Scanpy/AnnData (h5ad outputs).CITE-seq features mapped using TotalSeq C antibody reference; UMI counting performed as per 10x Feature Barcode pipeline.Human GRCh38 (hg38) for GEX; 10x vdj_GRCh38 human reference for V(D)J.CellBender-filtered 10x HDF5 matrices per capture (e.g., cellbender_output_NTX_1_filtered.h5).Integrated single-cell objects in AnnData format (.h5ad) with per-cell metadata (e.g., processed_adata_NTX_all_samples.h5ad), and per-cell V(D)J tables (per_cell_TCR_table.csv; my_merged_contigs.tsv). The uploaded files contain the environments, and base files for the analyses described in the M&M section of our Manuscript Regulon analysis with pySCENICGene‑regulatory networks (GRNs) and regulon activities were inferred with pySCENIC v0.11.2 (64). For each analyzed subset the raw count matrix was exported, all genes were flagged as highly variable, and human transcription factors (allTFs_hg38.txt) were retained. GRNs were inferred with GRNBoost2 (Arboreto v0.1.7) on a transient Dask v2024.3.0 cluster (12 workers, one thread each), producing TF‑target adjacency tables. Co‑expression modules (≥ 10 genes) were identified with modules_from_adjacencies and pruned for direct TF binding with prune2df against the hg38 Motif Collection v10 (10 kb up‑/down‑stream). Motif‑supported TF–target sets constitute the regulons. Per‑cell regulon activity was scored with AUCell (20 000‑gene ranking, default thresholds). Mean AUCell per subset ranked regulon activity, and the number of target genes served as a proxy for regulon size. Betweenness centrality network analysisTo quantify run‑to‑run stability, we repeated the pySCENIC workflow 20 independent times on the CAR‑CD4 subset defined by their respective barcodes. pySCENIC settings (GRNBoost2 inference, module detection, motif‑based pruning, AUCell scoring) were performed as described in the regulon analysis with pySCENIC section. Per run we saved adjacencies_CAR_CD4.csv, pruned_motifs_CAR_CD4.csv, regulons_CAR_CD4.pkl, and auc_mtx_CAR_CD4.csv, plus summary tables of regulon size and mean AUCell. These per‑iteration outputs were the inputs for cross‑run network aggregation and the panel below.For each SCENIC_CAR run we built a directed TF→target network from SCENIC outputs, preferring adjacencies_CELL.csv (columns: TF, target, importance) and falling back to regulons_CELL.pkl (extracting gene2weight). To avoid sparsity while prioritizing strong interactions, we kept the top 50 targets per TF and selected a data‑driven weight cutoff that yielded 1,500–20,000 edges per run. Degree pruning was disabled (LOW_DEG=0). No weight normalization across runs was applied (WEIGHT_TRANSFORM=None).A node was counted as present in a run if it appears in that run’s densified graph (i.e., it participates in ≥1 edge, in‑ or out‑going). Presence counts were accumulated over 20 independent SCENIC_CAR runs. We defined an overlap‑eligible node set as all genes present in ≥ ⌈0.90 × n_runs⌉ runs. The node‑overlap graph was then formed by aggregating all edges observed in any run between two overlap‑eligible nodes. For each edge (u→v), the across‑run weight equals the mean edge weight over runs where the edge exists, multiplied by its presence fraction (runs with the edge ÷ n_runs; SCALE_WEIGHTS_BY_PRESENCE = True, exponent = 1.0), which down‑weights run‑specific edges while retaining recurrent ones. The plotted graph includes all components (no LCC restriction). Node size encodes weighted betweenness centrality computed on the undirected projection with edge distance = 1/weight. We labelled the top 50 nodes by betweenness. Communities are colored by label propagation (undirected). Layout uses a spring embedder with fixed seed (SPRING_SEED = 4,572,321); random operations use RNG_SEED = 42. Citations64. Van de Sande B, Flerin C, Davie K, De Waegeneer M, Hulselmans G, Aibar S, Seurinck R, Saelens W, Cannoodt R, Rouchon Q, Verbeiren T, De Maeyer D, Reumers J, Saeys Y, Aerts S. A scalable SCENIC workflow for single-cell gene regulatory network analysis. Nat Protoc. 2020;15(7):2247-76.



