Supplementary Dataset and Benchmark Logs: Energy-Efficient SIMD-Accelerated Expression Parsing
收藏资源简介:
This repository contains the supplementary materials and experimental data supporting the research article: "Energy-Efficient SIMD-Accelerated Expression Parsing: A Cross-Architectural Microarchitectural Analysis in Managed Environments". The dataset is divided into three primary components: the benchmark logs and power telemetry, the evaluated ablation methodologies, and a summary of the key empirical findings across modern AMD Zen 4 and Intel Tiger Lake platforms. 1. Experimental Benchmarks and Hardware Telemetry This section contains the execution logs produced with BenchmarkDotNet v0.15.8 on .NET 10.0.12 (Windows 11). Unless stated otherwise, each case was measured 30 times (10 iterations across 3 process launches); the complexity-multiplier, code-size and ISA files use the default BenchmarkDotNet job. Files amd_zen4_baseline.md and intel_tigerlake_baseline.md: Latency and allocation for all pipeline stages and full-pipeline configurations, together with established .NET evaluators (System.Linq.Expressions, NCalc, Jace.NET, DataTable.Compute). Files amd_zen4_amortisation.md and intel_tigerlake_amortisation.md: Total cost as a function of the number of evaluations of one expression, compared with compiling it once to a delegate. Files amd_zen4_length_sweep.md, amd_zen4_depth_sweep.md and their Intel counterparts: Input length from 40 to 32,774 characters, and nesting depth from 1 to 15 in left- and right-associated forms. Files amd_zen4_error_path.md and intel_tigerlake_error_path.md: Rejection latency for twelve malformed inputs. Files amd_zen4_vector_width.md and intel_tigerlake_vector_width.md: Instruction-level latency, reciprocal throughput and per-element cost at 128, 256 and 512 bits. Files amd_zen4_isa_avx512_on.md and amd_zen4_isa_avx512_off.md: Scalar and Vector256 evaluation with and without AVX-512 code generation (DOTNET_EnableAVX512=0). File amd_zen4_evalonly_codesize.md: Evaluation stage with emitted JIT code size. Files amd_zen4_uprof_forward.csv, amd_zen4_uprof_reverse.csv and amd_zen4_uprof_shuffled.csv: AMD uProf traces (100 ms sampling) of per-core power and effective frequency for the sustained-power session in three segment orderings. File amd_zen4_uprof_dutycycle.csv: AMD uProf trace of the fixed 10 ms duty-cycle run used to measure active-state residency. Files amd_zen4_uprof_forward_markers.csv, amd_zen4_uprof_reverse_markers.csv, amd_zen4_uprof_shuffled_markers.csv and amd_zen4_uprof_dutycycle_markers.csv: Begin and end markers of every segment and idle window, emitted by the benchmark harness, against which the corresponding uProf trace is cut. Marker times are in UTC, whereas uProf timestamps are in local time (UTC+3). Files corpus_workload_manifest.csv and corpus_ulp_accuracy.csv: The 69 test expressions, and the scalar, Vector256 and Vector512 results for each with ULP difference and lane spread. File amd_zen4_complexity_multiplier.md: Latency and allocation of the pipeline stages under structural complexity multipliers of 10x and 100x. Its hardware-counter columns are not used in the article. 2. Evaluated Methodologies (Ablation Study) The benchmark data systematically isolates distinct parsing stages to test the following execution paths: ParseOnly: Character-by-character scalar tokenization (ScalarParse) against branchless AVX-2 classification (SimdParse). ConvertOnly: Stack-based Shunting Yard conversion (ShuntingYardConvert) against flat-array scanning (ScanBasedConvert). EvalOnly: Postfix evaluation with scalar, 256-bit (Vector256Eval) and 512-bit (Vector512Eval) arithmetic. FullPipeline: The full factorial matrix of tokenizer, converter and evaluator, with DataTable.Compute as a control. 3. Key Findings Highlighted in the Data Vector Width: On both platforms 512-bit operations keep the latency of 256-bit ones but issue at half the rate, so the 512-bit evaluator is slower (5.7% on AMD, 1.0% on Intel). Power and Energy (AMD only): The 512-bit path draws less dynamic core power than the 256-bit path. With fully occupied lanes and independent operations it consumes 23.8% less energy per element. Allocation: The evaluation stage allocates no managed memory. The complete pipeline allocates 1,216 B per operation, against 13,368 B for DataTable.Compute. Race-to-Sleep: Under a fixed duty cycle delivering identical work, active-state residency falls from 69.1% to 51.6%. Thresholds: The pipeline is faster than delegate compilation below about 230 evaluations on AMD and 310 on Intel, and vectorized tokenization loses its advantage beyond about 30,000 characters. All other hardware performance counter data from version 1 have been withdrawn, as explained in Section 4.4 of the article.



