aditijc/dhf-smoke-canary
收藏资源简介:
dhf-smoke-canary数据集是一个用于验证新的可解释性表面的数据集,包含16行数据和13个列。数据集的主要列包括代码ID、原始词汇字符串、带单位的英文标签、测量单位、屏蔽令牌数量、top1和top5预测计数及其准确率、预测和真实值的平均值(以临床单位表示)、均方误差以及超出范围的预测百分比。数据集生成参数详细说明了模型、实验名称、集群、状态、值统计路径、剪辑阈值路径、分割、批次、批处理大小以及消融摘要和指标。
The dhf-smoke-canary dataset is used to validate new interpretability surfaces, containing 16 rows of data and 13 columns. The main columns include code ID, raw vocabulary string, English label with units, measurement units, number of masked tokens, top1 and top5 prediction counts and their accuracy, mean predicted and true values (in clinical units), mean squared error, and percentage of predictions out of range. The dataset generation parameters detail the model, experiment name, cluster, status, value statistics path, clip thresholds path, split, batches, batch size, and ablation summaries and metrics.



