slopometer measures prose against a range of reference-prose rules. The rules merge GOV.UK/GDS house style with the discipline of ASD-STE100, including 26 numbered tells, banned-word lists, and a reference register. It also has rules covering a range of best practices for clear technical communication, including paragraph length, referant tracking, and restatement.
slopometer is for anyone shipping READMEs, docstrings, PR text, or agent responses. It catches slop mechanically, in milliseconds, with no LLM in the loop. Every finding cites a rule, quotes a span, and carries a weight.
pip install slopometer
The first score on a machine downloads spaCy’s en_core_web_md model (about 40MB, once) into ~/.cache/slopometer and loads it by path. The model never enters a virtual environment. One copy serves every project. Environment syncs cannot remove it.
from slopometer.score import score_text, score_pathThe default minimum is 150 scored words. Shorter inputs return too short to meter with their word count and no numeric score. Set min_words=0 to meter short examples:
score_text("This section describes our approach. It isn't just a linter - it's a comprehensive paradigm for quality.", min_words=0)density 194.1 (weight 33 on 17 prose words), worst 10
1: [10] notxbuty (tell 16, not-X-but-Y): "isn't just a"
1: [10] banned: 'comprehensive' -> 'complete'
1: [10] banned: 'paradigm'
1: [3] throat_clearing (tell 13, throat-clearing): 'This section describes'
Density is weighted findings per 100 words across paragraphs, headings, and list items. Heading and list markers do not count as words. Code blocks and other content excluded from scoring do not enter the denominator or count toward the minimum.
Each em dash adds 4 to the total weight. Semicolons joining independent clauses also add 4. Single and double hyphens add no weight on their own. Semicolons in lists of noun phrases or inside parentheses do not score. A comma and “and” before a clause with its own subject add 3. A sentence with three or more sections adds 1 for each section beyond two. A section ends at a comma, semicolon, colon, dash, spaced hyphen, or opening parenthesis. List items stay in one section.
score_text, score_path, score_paths, and score_many accept min_words, defaulting to 150. A short result has too_short=True, an empty findings list, and None for density, total, and worst.
score_path scores Markdown files, the Markdown cells of an .ipynb notebook, or the module docstring of a .py file. Notebook cells are joined with blank lines and scored as one document; code cells, outputs, and raw cells are excluded. File reports carry lineno|hash| addresses, prefixed with the cell ID for notebooks. For a .py file, addresses name lines of that file. Result.location(finding) returns the source location. score_paths scores several files and directories in parallel, one result per file. A directory contributes its .md, .qmd, .ipynb and .py files. The command line accepts the same paths, or Markdown on stdin:
slopometer README.md
slopometer nbs/00_core.ipynb
slopometer pkg/module.py
slopometer nbs/ README.md
git log -1 --format=%B | slopometer --min-words 0
slopometer draft.md --threshold 10
slopometer draft.md --pangram
--threshold returns exit code 1 when any file’s score exceeds the limit. Several paths or a directory give one report per file, highest density first. JSON output is a list with one entry per file. Each entry includes its path, which is null for stdin. Inputs below --min-words print the short-input message and exit successfully without scoring. JSON output marks them with too_short: true, includes min_words, and uses null scores. Every JSON finding includes a location object with one-based line, address, end_line and end_address. The end fields locate the finding’s last character. Notebook locations add cell_id and zero-based cell_index. Finding start and end offsets refer to the combined Markdown, not notebook JSON.
--pangram, pangram_text, pangram_path and pangram_paths score with Pangram’s AI detector instead of the rules. Each finding is a window of text that Pangram scored. Density runs from 0 (human) to 100 (AI). --threshold compares against the same scale. Each request sends the text to Pangram’s servers and costs about $0.05 per 100 words. The API key comes from PANGRAM_API_KEY.
The command runs warm through warmpy. The first input long enough to score loads the model in a background process. Later calls answer in milliseconds. After thirty idle minutes the process exits.
The rules live in notebooks that teach each family beside its code. Each rule states its tell, shows a violating example, and shows the plain rewrite. The lexicon notebook holds the word and phrase rules. The syntax notebook builds sentence rules on spaCy’s parse. The para notebook covers paragraph and document rules and the limits of word-vector heuristics. The score notebook assembles the pipeline. A drift test asserts that every write_docs tell maps to a rule or to an explicit unscoreable registry. The meter and the style guide cannot drift apart silently.
Vocabulary novelty is not scored. Identifying unexplained terminology requires audience and document context.
A rule ships only when its false-positive rate on clean reference prose is near zero. The meter scores the style guide’s own clean passage at exactly 0.0. The score notebook measures the blind spot instead of hiding it: mechanically chopped prose passes every surface rule while staying opaque. Agent review (check_docs) and Pangram scoring (--pangram) cover that residue. Vale, write-good, and proselint solve neighboring problems. The lexicon notebook records what came from them. It also credits the GOV.UK words-to-avoid list (OGL v3).