primus.scoring package

Programmatic scoring entry point for eval pipelines.

This module exposes a thin, dependency-light scoring API that eval pipelines (scanner corpus eval, CI gates) can call directly without starting a GraphQL server, seeding scorecard metadata, or importing the dashboard/Evaluation orchestration stack.

The scoring logic reuses the same span-overlap matching as primus.scores.SourceSpanOverlapScore but is invoked as a plain function over in-memory data, not as a Score class over Items.

primus.scoring.evaluate_recall(annotations: list, findings: list) → dict[str, Any]

Score recall of findings against gold annotations.

Each annotation is a dict with file_path, start_line, end_line, and expected.status ("positive" or "negative"). Each finding is a dict with filePath, startLine, endLine.

Returns a dict with accuracy, recall, precision, confusion_matrix, and denominators. Negative annotations are scored as “No” references; a scanner finding overlapping them is a false positive.

Submodules