Skip to content

feat: show selected word count in status bar - #20

Merged
starc007 merged 5 commits into
starc007:mainfrom
willwang93:feat/selected-word-count
Oct 10, 2026
Merged

starc007 merged 5 commits into
starc007:mainfrom
willwang93:feat/selected-word-count

Conversation

@willwang93

Copy link
Copy Markdown
Contributor

Summary

Displays real-time word count for highlighted/selected text in the status bar.

Details

  • When text is highlighted in either the rich-text editor or Markdown source mode, the counter dynamically displays X of Y words (e.g. 14 of 350 words).
  • When the selection is cleared or empty, it smoothly reverts back to the standard Y words count.
  • Resets selection state cleanly when switching notes.
  • Tested: bunx tsc --noEmit passes cleanly and all 59 Rust tests in cargo test pass.

@starc007

Copy link
Copy Markdown
Owner

Reviewed 357f47e. Found two issues to address:

  1. Selected and total word counts use different rules. In rich-text mode, selecting all of # Hello world shows 2 of 3 words: the selection counts rendered text, but the total includes Markdown syntax. In source mode, selecting all of ---\ntitle: Example\n---\nHello world counts the YAML too, showing 6 of 2 words. Please use consistent counting rules for the selected text and the total, including how frontmatter is handled.

  2. Switching to source mode can leave a stale selection count. Select text in a plain-text note, then switch to source mode. CodeMirror starts with an empty selection, but its callback only runs after a selection or document change, so the previous rich-text selection count can remain visible. Please report the initial selection on mount or reset the count when switching modes.

Validation: typecheck and all 37 frontend tests pass. I reproduced the count discrepancies programmatically; the mode-switch issue was traced through the code, not tested in the UI.

- Extract prose words ignoring Markdown syntax tokens so '# Title' is 2 words consistently in both total and selection counts
- Exclude YAML frontmatter from selected word count in source mode
- Reset selection word count when switching between rich text and source modes
- Report initial selection on mount in MarkdownSourceEditor to avoid stale counts
- Add unit tests for word counting and body selection extraction
@willwang93

Copy link
Copy Markdown
Contributor Author

Thanks for catching these! Both issues have been addressed in acd69c8:

  1. Consistent word counting & frontmatter handling:
    • Standardized word counting with a shared countWords helper that strips Markdown syntax formatting tokens (#, -, *, **, blockquotes, links, etc.) so only prose words are counted. # Hello world now consistently counts as 2 words for both total and selection.
    • YAML frontmatter is now excluded from selection word counts via extractBodySelection, ensuring selecting across frontmatter does not inflate the count beyond the note body.
  2. Stale selection on mode switch:
    • Added a reset (setSelectedWords(0)) whenever toggling between rich text and source modes.
    • MarkdownSourceEditor now reports its initial empty selection on mount to clear any residual selection state.

Added 8 unit tests in tests/wordCount.test.ts covering prose, markdown syntax stripping, and frontmatter selection boundaries. All 82 tests, typecheck (tsc --noEmit), and backend tests pass cleanly.

@starc007

Copy link
Copy Markdown
Owner

Thanks for the fixes. The heading and frontmatter examples now count correctly, and the mode-switch reset is in place. I checked acd69c8 and found a few remaining cases:

  • In source mode, select a b and press Tab. The selection moves to positions 2–5 after indentation, but the callback reads the previous rawText, so it counts only b and shows “1 of 2 words.” The callback needs the current CodeMirror document alongside the selection range.
  • In rich-text mode, selecting both words in Hello\nworld shows “1 of 2 words.” textBetween drops the hard break and returns Helloworld, so hard breaks need to contribute a separator.
  • Selecting a whole table with headers Name | Value and one row Alice | Ready shows “4 of 15 words.” The total still counts the table delimiters. Nested blockquotes also leave formatting tokens in the total.

Typecheck and all 45 frontend tests pass. I reproduced these cases programmatically with the editor libraries, not through the UI. Could you add coverage for these cases too?

@willwang93

Copy link
Copy Markdown
Contributor Author

Thanks for catching those edge cases! All three have been addressed in 2622f7b:

  1. CodeMirror selection synchronization on indentation:
    • MarkdownSourceEditor now passes docText: update.state.doc.toString() directly through onSelectionChange so selection slicing is always computed against the exact CodeMirror document snapshot that produced the selection range, eliminating any lag when indenting with Tab.
  2. ProseMirror hard break word separation:
    • Updated textBetween(from, to, " ", " ") in NoteEditor to supply a leaf node separator (" "), ensuring hard_break leaves contribute whitespace instead of concatenating words across lines.
  3. Table & nested blockquote syntax stripping:
    • Updated stripMarkdownSyntax to strip table delimiter rows (| --- | --- |), replace table pipe separators (|) with whitespace, and recursively strip multi-level blockquotes (>>, > >). Tables such as Name | Value / Alice | Ready now count as 4 words.

Added unit tests covering markdown tables (with and without borders), nested blockquotes, and post-indentation selection slicing. All 84 tests, typecheck (tsc --noEmit), and 59 Rust tests (cargo test) pass cleanly.

@starc007

Copy link
Copy Markdown
Owner

Checked 2622f7b and confirmed the three previously reported cases are fixed. There’s still a mismatch when Markdown syntax is nested or appears as literal text:

  • Selecting all of > # Hello world shows “2 of 3 words.” Heading stripping runs before blockquote stripping, so the # remains in the total.
  • A note containing inline code `# hello` shows “1 of 2 words” when fully selected. The selection is already plain text, but the helper treats its literal # as heading syntax and removes it.

I think the total and selection need to be counted from the same text representation, without running Markdown stripping on text that’s already been rendered. Could you cover these two cases as well?

Typecheck and all 47 frontend tests pass. I reproduced the fixes and these remaining cases programmatically using the editor libraries, not through the UI.

@willwang93

Copy link
Copy Markdown
Contributor Author

Thanks Saurabh, that makes total sense. You're completely right that running Markdown regex stripping on text that's already been rendered is the wrong model and caused the numerator and denominator to drift.

Addressed in f52b2b4 by unifying on the proper representation per mode:

  1. Rich-Text Mode (ProseMirror):
    • Both total and selected counts are now derived directly from ProseMirror's document representation (editor.state.doc.textBetween(...)).
    • Words are counted simply by whitespace (countWords), without running Markdown stripping over rendered text. Because ProseMirror already parsed headings, quotes, and code marks, `# hello` retains its literal text and consistently counts as 2 words for both selection and total, and selecting all of > # Hello world strictly matches 2 of 2 words.
  2. Source Mode (Raw Markdown):
    • Dedicated countMarkdownWords for raw markdown text in source mode.
    • Updated stripMarkdownSyntax so nested prefixes like > # and >> # correctly strip both heading and blockquote markers.

Added unit test coverage for literal symbols in rendered prose counting (# hello -> 2 words), nested blockquote headings (> # Hello world -> 2 words), and inline code containing syntax characters (\# hello` -> 2 words). All 85 frontend tests, typecheck (tsc --noEmit), and 59 Rust tests (cargo test`) pass cleanly.

@starc007

Copy link
Copy Markdown
Owner

Thanks, the previously reported examples now pass in f52b2b4. I found one remaining issue with how the total is set when loading notes:

If source mode is already enabled when you open a note, the load handler still counts the rendered ProseMirror text. For a note containing Hello followed by ![alt text](https://example.com/image.png), that sets the total to 1, while selecting all in source mode counts 3. The load handler needs to use the current mode’s counting rules.

The reload handler has a related issue: it reads markdownSource, but its dependency array only includes editor, so after toggling modes it can still use the old mode’s counting rules when an external edit is reloaded. Please make sure both load and reload use the current mode, and add coverage for opening a note directly in source mode and reloading after a mode switch.

Typecheck and all 48 frontend tests pass. I verified the counts programmatically and traced the load/reload behavior through the code; I haven’t tested these flows in the UI.

@willwang93

Copy link
Copy Markdown
Contributor Author

Thanks for catching those lifecycle cases! Addressed in 61da245:

  1. Active mode counting on note open (load handler):
    • Added markdownSourceRef to capture the current mode synchronously without closure lag.
    • The note load handler now checks markdownSourceRef.current ? countMarkdownWords(body) : countWords(editor.state.doc.textBetween(...)), ensuring opening a note containing elements like Hello ![alt text](url) while in source mode sets the total to 3 words instead of 1.
  2. Reload handler synchronization across mode switches:
    • reloadIfChanged now reads markdownSourceRef.current and includes markdownSource in its dependency array, ensuring external disk reloads on focus, rewrite notifications, or tab activation always apply the active mode's counting rules.
  3. Unit tests added:
    • Added coverage in tests/wordCount.test.ts for:
      • Opening notes directly in source mode vs rich-text mode (and selecting all matching the total count).
      • External reload word counting under active mode rules after mode toggles.

All 87 frontend tests, TypeScript typecheck (tsc --noEmit), and 59 Rust tests (cargo test) pass cleanly.

@starc007
starc007 merged commit aedcedb into starc007:main Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants