Fast, lossless LLM inference via dual-view diffusion decoding.
-
Updated
May 18, 2026 - Python
Fast, lossless LLM inference via dual-view diffusion decoding.
Official Implementation of "Gumiho: A Hybrid Architecture to Prioritize Early Tokens in Speculative Decoding" (ICML'25)
Biological code organization system with 1,029+ production-ready snippets - 95% token reduction for Claude/GPT with AI-powered discovery & offline packs
Reduce Claude AI token consumption. Zero install to start. Python 3.7+ for auto-manifest generation.
Too lazy to waste tokens. Save input, output & media tokens on Claude, ChatGPT, Gemini.
Advanced token reduction and prompt optimization framework for LLMs, featuring linguistic, algorithmic, and architectural patterns.
Dynamic latent-state control heads for LLMs: route each query by actual model capability, not task type, to answer, reason, call tools, abstain, or escalate
Professional AI token optimization software that reduces token usage by up to 75% for developers and teams. Boost efficiency and lower AI expenses in 2026.
Packet-Switched Attention for stable 2-bit quantized MoE inference, with variance-aware routing and Protocol C benchmarks.
Do dense LMs develop MoE-like specialization as they scale? Measure it, visualize it, and turn it into speed.
Powerful AI efficiency tool that reduces token usage by up to 75% for cloud code and LLM applications. Ideal for developers looking to maximize performance while minimizing costs in 2026.
BitNetLab — Interactive 1-bit / 1.58-bit (ternary) LLM Quantization & Matmul-Free Inference Laboratory. Live QAT with the straight-through estimator, BitLinear, bit-width vs quality, memory/energy, activation outliers.
TokenCave is a browser extension for Claude AI that helps you monitor and optimize token usage with real-time counters, usage insights, and a “caveman mode” that dramatically reduces output length while preserving technical accuracy.
TIDAL: Toolkit for Inference, Deployment, Adaptation, and Learning for efficient AI systems
Add a description, image, and links to the llm-efficiency topic page so that developers can more easily learn about it.
To associate your repository with the llm-efficiency topic, visit your repo's landing page and select "manage topics."