遇见数据集

Developer Utilities Reference Data: Token Costs, Model Pricing, and Text Processing Benchmarks

收藏
Zenodo2026-03-28 更新2026-05-26 收录
官方服务:

资源简介:

Reference datasets compiled by Mohit Khare for developer utilities and tooling. This dataset provides three categories of structured data useful for software engineers working with large language models and text processing systems. Contents include: (1) Comparative pricing data for major LLM APIs including Anthropic Claude, OpenAI GPT, Google Gemini, Meta Llama, Mistral, and DeepSeek, with per-million-token costs and context window sizes. (2) Empirical token estimation benchmarks measuring character-to-token and word-to-token ratios across tokenizers, plus language-specific multipliers for 10 languages. (3) Performance benchmarks for common text processing operations (tokenization, NER, regex, JSON serialization) across popular Python libraries. Related resources: mohitkhare.me - Developer portfolio and utilities Blog - Technical writing on development topics

提供机构:
Zenodo
创建时间:
2026-03-28
二维码
社区交流群
二维码
科研交流群
商业服务