forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 26
Pull requests: unslothai/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Prebuilt: move the #95 pin to the identity-guard commit
#98
opened Aug 11, 2026 by
danielhanchen
Member
Loading…
sampling: index penalties by token id instead of scanning every candidate
#95
opened Aug 11, 2026 by
danielhanchen
Member
Loading…
CI: let a cold CUDA profile warm its ccache from a sibling profile
#83
opened Aug 8, 2026 by
danielhanchen
Member
•
Draft
3 tasks
Prototype: pin the CPU legs' glibc floor with a container, not the runner
#79
opened Aug 7, 2026 by
danielhanchen
Member
•
Draft
kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes
#70
opened Aug 6, 2026 by
danielhanchen
Member
Loading…
IQ1_XS, IQ1_XXS, IQ1_XXXS: three quant types below IQ1_S
#61
opened Aug 3, 2026 by
danielhanchen
Member
Loading…
Prebuilt: drop the MiniMax-M3 and Inkling pins from the PR set
#43
opened Jul 28, 2026 by
oobabooga
Member
Loading…
metal: fall back to copied weight buffers on macOS 14 paravirtual devices
#37
opened Jul 22, 2026 by
danielhanchen
Member
Loading…
mistral3: diagnosis of long-context degradation on Mistral-Medium-3.5-128B
#16
opened May 1, 2026 by
danielhanchen
Member
Loading…
2 tasks
WIP: DeepSeek-V4-Flash architecture port (BF16 reference, accuracy-first MVP)
#14
opened Apr 29, 2026 by
danielhanchen
Member
•
Draft
1 of 9 tasks
ProTip!
Exclude everything labeled
bug with -label:bug.