P8-4: skip the per-function source re-slice below the copy floor

`duplicate_bodies` extracted a source segment for EVERY function in the module
population, including one-liners that can never reach the ten-line literal-copy
floor. `ast.get_source_segment` re-splits the whole file per call, so the scan
paid O(source) for each short function.

The span a node already carries (`end_lineno - lineno + 1`) decides whether the
slice is worth taking. This is output-identical by construction, not merely by
test: `_normalize_body` dedents, rstrips and DROPS blank lines, so a normalized
body can never have more lines than its raw span, and a node whose span is below
the floor could never have satisfied the existing normalized-line check. The
recursion into nested functions stays unconditional, so a long inner function
inside a short outer one is still scanned.

Oracle, over the full module population with every module given its own domain
so every digest group surfaces as a row: byte-identical output (3 rows, same
sha256) before and after, 31.1s -> 21.4s. `scripts/check_domains.py` exits 0
with "OK: domain manifest complete", and the `[duplicates] allowed = []`
baseline is untouched in both directions.

The new boundary test pins the floor exactly: a cross-domain copy of exactly
DUPLICATE_MIN_LINES lines is still detected, one line shorter is not. Verified
load-bearing - changing the filter to `span > min_lines` reddens it.
This commit is contained in:
Anton 2026-09-11 23:25:45 +03:00 • committed by Ouroboros
parent 1fbc818c44
commit 065467f97d
2 changed files with 26 additions and 1 deletions

View file

@ -403,7 +403,10 @@ def duplicate_bodies(manifest: Manifest, root: pathlib.Path = REPO_ROOT,
for child in ast.iter_child_nodes(node):
if isinstance(child, (ast.FunctionDef, ast.AsyncFunctionDef)):
qual = f"{prefix}{child.name}"
segment = ast.get_source_segment(source, child)
# A normalized body is never longer than its raw span (blank
# lines are dropped), so a span below the floor cannot reach it.
span = (getattr(child, "end_lineno", None) or child.lineno) - child.lineno + 1
segment = ast.get_source_segment(source, child) if span >= min_lines else None
if segment is not None:
norm = _normalize_body(segment)
if norm.count("\n") + 1 >= min_lines: