Popular repositories Loading
-
chronos-forecasting
chronos-forecasting PublicChronos: Pretrained Models for Time Series Forecasting
-
-
Repositories
- hallucination-benchmark-trivialplus Public
[ACL 2026 main] Long-Context Hallucination Detection Benchmark: Rethinking Evaluation for LLM Hallucination Detection: A Desiderata, A New RAG-based Benchmark, New Insights
- carbon-assessment-with-ml Public
CaML: Carbon Footprinting of Household Products with Zero-Shot Semantic Text Similarity
- SWE-PolyBench Public
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
- StaminaBench Public
- reskill Public
An easy-to-configure and extensible veRL extension for agent RL training with skill co-evolution.
- muss Public
MUSS: Multilevel Subset Selection for relevance and diversity at scale (UAI 2026) — up to 80x faster than MMR with approximation guarantees, for RAG and candidate retrieval
- tabpfn-automl2026 Public
- gpbm Public
- Self-Evolving-Agents-Double-Ratchet Public
Reference implementation of Double Ratchet: co-evolving an inspectable evaluation metric with a lifecycle-managed skill library for self-improving LLM agents (arXiv:2607.12790)
- Self-Evolving-Agents-Ratchet Public
Reference implementation of Ratchet: a minimal hygiene recipe for self-evolving LLM agents (arXiv:2605.22148)
Most used topics
Loading…