Best Token Use Path (BTUP)
Plain-language summary of a PacketViper preprint. DOI 10.5281/zenodo.21337900, published July 13, 2026.
Last reviewed: October 2026. This page is a plain-language summary built only from the paper’s own abstract and text. It is not a substitute for the paper; read the full text on Zenodo.
At a glance
| Item | Detail |
|---|---|
| Title | Best Token Use Path (BTUP): Client-Side Traffic Engineering for Cost-Aware LLM Inference |
| Author | Francesco Trama (ORCID 0009-0004-8437-6351) |
| Published | July 13, 2026 (Zenodo preprint) |
| Version | Not stated on the Zenodo record (file: v1.0 release candidate) |
| DOI | 10.5281/zenodo.21337900 |
| License | Creative Commons Attribution 4.0 (CC BY 4.0) |
| Full text | Read on Zenodo |
Plain-language summary
Large language models are trained toward a quality objective that contains no representation of what a response costs in tokens. The paper compares this to a link-state routing protocol whose advertisements omit the cost field: every path looks free, so the system tends to prefer verbose, expensive routes such as preambles, restatements, hedging, full-file rewrites and unbounded agent loops. Earlier work addresses pieces of this problem, including model cascades and routing, prompt compression, reasoning budgets and length control, but usually as isolated optimizations.
The paper introduces Best Token Use Path (BTUP), a provider-independent, client-side traffic-engineering framework. BTUP treats an AI request as a complete execution path: context policy, retrieval, prompt transform, cache strategy, provider and model, reasoning and output budgets, output contract, tool plan, validation, and retry or escalation policy. It selects the path that minimizes expected cost per accepted resolution, subject to quality, latency, safety, privacy and budget constraints. The author specifies the control plane, a provider-accounting reconciliation model, privacy-preserving telemetry, and a security model covering denial-of-wallet, tool-loop containment and refusal handling.
The author argues the approach grows in importance regardless of price direction, because agentic workloads multiply tokens per task faster than unit prices decline. In the version reviewed for this summary, the production-results sections are marked as pending, so this summary reports the framework and its hypotheses, not measured results. This is an AI-systems research paper, not a PacketViper product paper.
What the paper does not claim
How to cite
Trama, F. (2026). Best Token Use Path (BTUP): Client-Side Traffic Engineering for Cost-Aware LLM Inference. Zenodo preprint, Not stated on the Zenodo record (file: v1.0 release candidate). https://doi.org/10.5281/zenodo.21337900
Explore further
The paper, the author and related pages.