fix(consolidate): keep whole words that fit when a rule slug hits the cap

Follow-up to #886. The word-boundary cut searched `out[..60]` for a
hyphen, so a slug whose first 60 characters ended exactly on a word
(hyphen at index 60) dropped that word. And any hyphen counted, however
early: `a` followed by one 70-letter token collapsed the slug to `a`,
losing the rest of the title.

The window now includes index 60, and a hyphen in the first half is
ignored, so the cut falls back to the hard 60. Two regression tests;
reverting either change turns exactly its own test red.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01K7RSKQrakhnAfKjGaNaVR2
This commit is contained in:
Samir Hanna Verza
2026-09-25 10:29:40 -03:00
co-authored by Claude Opus 5.5
parent 8629adc665
commit 4d7d169071
2 changed files with 33 additions and 1 deletions
+7
View File
@@ -38,6 +38,13 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0
feedback-loop guard stays non-overridable. (#894)
### Fixed
- Rule slugs that hit the 60-character cap keep every whole word that fits.
The word-boundary cut from #886 only looked for a hyphen before position 60,
so a slug whose first 60 characters ended exactly on a word dropped that
word, and a hyphen early in the title (a short first word before one long
token) collapsed the slug to that single word. The cut now counts a hyphen
at position 60 and ignores one in the first half, falling back to the hard
cut at 60. (#886)
- A manual `memory_consolidate` now reconciles the session's durable
consolidation job row. The MCP handler wrote the page directly through the
consolidator without touching `session_consolidation_jobs`, so a session
@@ -1544,7 +1544,11 @@ fn slugify_for_rule(title: &str) -> String {
// `out` is ASCII here, so byte index 60 is a char boundary. Cut at
// the last hyphen inside the window to end on a whole word; only
// hard-cut at 60 when the window holds no hyphen (one long token).
match out[..60].rfind('-') {
// The window includes index 60: a hyphen there means the first 60
// chars are whole words, and they all fit. A hyphen in the first
// half does not count, because cutting there would throw most of
// the title away (a short first word before one long token).
match out[..=60].rfind('-').filter(|&idx| idx >= 30) {
Some(idx) => out.truncate(idx),
None => out.truncate(60),
}
@@ -1961,6 +1965,27 @@ mod tests {
assert_eq!(slugify_for_rule("中文标题"), "rule");
}
/// A slug whose first 60 chars already end on a whole word keeps that
/// word: the hyphen right after it (index 60) is the boundary, and a
/// window that stops before it dropped the word (follow-up to #886).
#[test]
fn slugify_keeps_a_word_that_ends_exactly_at_the_cap() {
let title = ["abcd"; 11].join(" ") + " abcde more";
let slug = slugify_for_rule(&title);
assert_eq!(slug, ["abcd"; 11].join("-") + "-abcde");
assert_eq!(slug.len(), 60);
}
/// A boundary in the first half would throw most of the title away: a
/// short word before one long token must not collapse the slug to that
/// word, so the cut falls back to the hard 60 (follow-up to #886).
#[test]
fn slugify_does_not_collapse_to_a_short_first_word() {
let slug = slugify_for_rule(&format!("a {}", "b".repeat(70)));
assert_eq!(slug.len(), 60);
assert!(slug.starts_with("a-bbb"), "slug collapsed to {slug:?}");
}
fn update_with_summary(summary: Option<&str>) -> crate::types::ConsolidatedPageUpdate {
crate::types::ConsolidatedPageUpdate {
path: "concepts/queue.md".into(),