遇见数据集

DCTR: Pythia e+e- -> Z -> dijets datasets

收藏
Zenodo2020-07-30 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

A collection of datasets used in Neural Networks for Full Phase-space Reweighting and Parameter Tuning. Sample code for reproducing the results is available on GitHub. Each dataset was generated with the Pythia 8.230 event generator. Particle-level \(e^+ e^- \to Z \to \text{dijet}\) events with about 100 particles in each event are clustered into jets using the anti-kt clustering algorithm (R = 0.8) with Fastjet 3.0.3. Every jet is presented as a list of constituents \((p_T, \eta, \phi, \text{particle ID}, \theta)\) where \(\theta = (\texttt{TimeShower:alphaSvalue}, \texttt{StringZ:aLund }, \texttt{StringFlav:probStoUD})\). <strong>Training Datasets:</strong><br> Each training file contains two arrays, X and Y.<br> X is an array of jets and Y is 0 (1) if the jet was generated with default (non-default) Pythia parameters. For a non-default jet (Y=1), the \(\theta\) in each constituent represents the value of the Pythia parameter that was used. Note that for a default jet (Y=0) the \(\theta\) in each constituent is <strong>not</strong> the default Pythia parameters, but \(\theta\) uniformly sampled in the same range as the Y=1 jets. The parameters were uniformly sampled in \(\texttt{TimeShower:alphaSvalue} \in [0.10, 0.18]\) \(\texttt{StringZ:aLund } \in [0.50, 0.90]\) \(\texttt{StringFlav:probStoUD } \in [0.10, 0.30]\) The 1D datasets are labeled by which parameter was changed, and the 3D dataset simultaneously vary all three parameters. <strong>Test Datasets:</strong><br> Each test dataset consists of an dictionary containing: 'jet': the jet constituents 'multiplicity': Number of particles in jet 'tau21': Nsubjettiness observable 'tau32': Nsubjettiness observable 'ECF_N3_B4': Energy Correlation Function(N=3, \(\beta\)=4) 'ECF_N4_B4': Energy Correlation Function(N=4, \(\beta\)=4) The corresponding \(\theta\) values for each test set are described in the paper.

本数据集集合面向全相空间重加权与参数调优的神经网络研究。复现研究结果的示例代码已发布于GitHub平台。所有数据集均通过Pythia 8.230事件生成器构建生成。粒子级别的正负电子湮灭至Z玻色子并碎裂为双喷注((e^+ e^- o Z o ext{dijet}))事件中,每个事件约包含100个粒子,通过Fastjet 3.0.3工具包的anti-kt聚类算法(参数R=0.8)聚类为喷注。每个喷注以其构成粒子的列表形式表示,列表元素为((p_T, eta, phi, ext{粒子ID}, heta)),其中( heta = ( exttt{TimeShower:alphaSvalue}, exttt{StringZ:aLund }, exttt{StringFlav:probStoUD}))。<strong>训练数据集:</strong><br>每个训练文件包含两个数组X与Y。X为喷注数组,若喷注采用默认Pythia参数生成,则Y=0;若采用非默认Pythia参数生成,则Y=1。对于非默认参数生成的喷注(Y=1),其每个构成粒子的( heta)代表所使用的Pythia参数值。需注意,对于默认参数生成的喷注(Y=0),其每个构成粒子的( heta)并非默认Pythia参数,而是与Y=1类喷注采样范围一致的均匀分布采样值。上述参数均采用均匀采样,采样范围为:( exttt{TimeShower:alphaSvalue} in [0.10, 0.18]),( exttt{StringZ:aLund} in [0.50, 0.90]),( exttt{StringFlav:probStoUD} in [0.10, 0.30])。一维数据集按照被修改的参数进行标注,三维数据集则同时调整全部三个参数。<strong>测试数据集:</strong><br>每个测试数据集由包含以下字段的字典构成:'jet':喷注的构成粒子;'multiplicity':喷注内的粒子数;'tau21':Nsubjettiness可观测量;'tau32':Nsubjettiness可观测量;'ECF_N3_B4':能量相关函数(N=3,(eta)=4);'ECF_N4_B4':能量相关函数(N=4,(eta)=4)。各测试集对应的( heta)参数值详见相关论文。

提供机构:
Zenodo
创建时间:
2019-10-25
二维码
社区交流群
二维码
科研交流群
商业服务