Add ScoringBench integration and harden release evidence - #19
Draft
jxucoder wants to merge 49 commits into
Draft
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
1027_ESLquality sentinel on Linux and freeze its manifest, dataset registry, and raw fold-level Parquet resultsAGENTS.mdplus alearnings/engineering logbatch_sizeoption is setWhy
OpenBoost needs externally reproducible evidence before it can credibly claim a niche in distributional tabular boosting. The existing benchmark and scaling documentation overstated several unverified paths, while categorical persistence and cardinality handling could silently change predictions.
This PR establishes a truthful benchmark path, fixes identified correctness bugs, records failures as well as successes, and makes unsupported scaling behavior explicit.
First real-dataset evidence
The frozen ScoringBench shard is in
benchmarks/evidence/scoringbench/1027_esl_20260816/.Protocol: ScoringBench commit
a938a667..., 5 folds, 3,000-row cap, OpenBoost NaturalBoost Normal CPU vs NGBoost Normal, 500 rounds, learning rate 0.01, depth 3, seed 42.Descriptive means:
This is one small dataset, not a library-level quality or speed claim.
Verification
9256929555, digestsha256:0c5176c4a28cd44444f6a086643a01f636b0dcb5e165a4a30d3d992f3e96da97Remaining evidence gate
Keep this PR draft. One sentinel is not sufficient. Before any broad value claim: