primus.scoring package
Programmatic scoring entry point for eval pipelines.
This module exposes a thin, dependency-light scoring API that eval pipelines (scanner corpus eval, CI gates) can call directly without starting a GraphQL server, seeding scorecard metadata, or importing the dashboard/Evaluation orchestration stack.
The scoring logic reuses the same span-overlap matching as
primus.scores.SourceSpanOverlapScore but is invoked as a plain function
over in-memory data, not as a Score class over Items.
- primus.scoring.evaluate_recall(annotations: list, findings: list) dict[str, Any]
Score recall of
findingsagainst goldannotations.Each annotation is a dict with
file_path,start_line,end_line, andexpected.status("positive"or"negative"). Each finding is a dict withfilePath,startLine,endLine.Returns a dict with
accuracy,recall,precision,confusion_matrix, and denominators. Negative annotations are scored as “No” references; a scanner finding overlapping them is a false positive.