treepeat is a tool that finds similarities in your codebase.
- Find duplicate code blocks meaningful to the language (classes/functions), not just lines (--ruleset none)
- Find near-duplicates, ignore whitespace, strings, high level AST nodes such as function and names (--ruleset default)
- Find structurally similar code, anonymizing identifiers, constants, etc (--ruleset loose)
Pull requests welcome: This is very much an proof of concept - I'm happy with it, but I haven't supported very many languages at present. PRs welcome!
Languages supported: astro, bash, css, go, html, javascript, lua, markdown, python, sql, typescript, java, kotlin, rust, yaml
pip install treepeatScan a codebase for similar or duplicate code blocks using tree-sitter AST analysis and locality-sensitive hashing.
Key flags:
- --ruleset: Normalization ruleset to use (- none,- default,- loose) - controls how code is normalized before comparison
- --similarity: Percent similarity from 1-100 (default: 100 for exact duplicates)
- --min-lines: Minimum number of lines for a match (default: 5)
- --diff: Show side-by-side comparisons of similar blocks
- --format: Output format -- console(default) or- sariffor CI integration
- --verbose: Show additional run metrics, including per-stage timing when available
- --progress: Show progress bars for long-running pipeline stages
# Find exact duplicates
treepeat detect /path/to/codebase
# Find near-duplicates with 80% similarity threshold
treepeat detect --similarity 80 /path/to/codebase
# Show diffs between similar blocks and use loose ruleset
treepeat --ruleset loose detect --diff --min-lines 10 /path/to/codebase
# Show parse progress and verbose stage timing
treepeat detect --progress --verbose /path/to/codebase
# Output results in SARIF format for CI tools
treepeat detect --format sarif -o results.sarif /path/to/codebase--progress is intended primarily as interactive CLI feedback. The current implementation writes tqdm progress bars to stderr, leaving normal command output on stdout or --output.
List all rules in a ruleset, along with their descriptions. Use --language to see which rules apply to a specific language.
Display how treepeat normalizes source code into tree-sitter tokens for similarity detection -- helpful for debugging why a certain section of a file might be similar to another. Shows the original source code side-by-side with the normalized token representation.
make setup
make testPerformance harness for benchmarking real repositories: tools/perf/harness.py (docs).
Architecture Decision Records live in docs/adr.