* fix(auto-reply): treat U+2028/U+2029 as paragraph boundaries when chunking
chunkByParagraph normalized only CR/CRLF before blank-line paragraph detection,
so model output using Unicode LINE/PARAGRAPH SEPARATOR (U+2028/U+2029) instead
of a blank line was not split at those boundaries and fell back to length-based
splitting. Normalize U+2028/U+2029 to \n alongside CR/CRLF, matching how the
Control UI markdown renderer handles them.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
* fix(auto-reply): normalize U+2028 as line break, U+2029 as blank-line paragraph boundary
Distinguish U+2028 (LINE SEPARATOR) from U+2029 (PARAGRAPH SEPARATOR):
U+2029 becomes \n\n (blank line — paragraph boundary) while U+2028
becomes \n (single newline — intra-paragraph line break).
The original fix mapped both to \n, so standalone U+2029 still
produced single-line text without a blank-line gap — paragraph
detection failed. The combined U+2028 input accidentally
produced the right blank-line sequence, which masked the bug.
Adds individual tests for lone U+2029 (splits at paragraph boundary),
lone U+2028 (stays within paragraph), and consecutive U+2028
(combined blank line — matches \n\n behavior).
* test(auto-reply): simplify Unicode separator cases
---------
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* test(utils): add unit tests for parseJsonWithJson5Fallback
* test(utils): verify JSON.parse fast path with spy, per ClawSweeper review
- Add vi.spyOn to prove JSON.parse is called (not JSON5) for strict JSON
- Add spy to prove JSON5.parse is called on fallback path
- Run full src/utils test suite (263 tests passed)
* feat(gateway): typed structured questions on openclaw.chat
The chat result gains an additive optional question field (2-4 options, one
recommended max, per-option reply text). Producers: the two onboarding welcome
variants (apply-setup yes/ask, ready next-step) and hosted wizard select/confirm
steps, which mirror the awaited step so card clients render real wizard choices.
Prose replies always stand alone for text-only clients (macOS app, TUI). The
custodian page consumes the typed field and drops the PR1 string-marker parser
(never shipped in a release); cards send the option reply while the transcript
shows the label.
* chore: keep OnboardingWelcome type module-local
* chore(protocol): regenerate Swift bindings for the chat question field
When decodeMcpAppSandboxCsp receives a base64-encoded value that
decodes to non-JSON text, JSON.parse throws an unhandled exception.
Add a try-catch so the function returns undefined instead of crashing.
* test(utils): add unit tests for formatTokenCount
* fix(test): use Number.NaN/Number.POSITIVE_INFINITY per oxlint unicorn/prefer-number-properties
* test(utils): fold formatTokenCount boundary tests into existing usage-format test
Move edge case coverage (invalid inputs, zero/negative, exact boundaries,
thousands overflow) into src/utils/usage-format.test.ts instead of a
separate colocated test file per reviewer feedback.
* perf(ci): cut the pre-fan-out critical path on canonical runs
Two changes to the run head and matrix shape:
1. runner-admission was a hosted 90s sleep every run queued behind
(observed 1.7min hosted-queue latency before the sleep started). The
debounce now lives at the tail of preflight: heavy jobs all need
preflight, so a superseding main push still cancels the run before
fan-out while only one 4 vCPU runner has been spent, and preflight's
own work usually exceeds the window so the residual sleep is zero.
security-fast (hosted, dependency-free) starts immediately.
2. Canonical main pushes now use the compact bin plan like PRs: the
82-job named matrix drained the runner pool for ~4.5min (job starts
trickled from minute 5.2 to 9.7 in run 29592647843) with no
branch-protection consumer for per-shard names on main. Coverage is
identical; dispatch/release-gate runs keep the full named matrix.
* perf(test): boot TUI PTY suite fixtures concurrently
tsx+TUI startup dominated the harness file's wall time and the three
suite PTYs booted serially. Boot them concurrently in beforeAll
(allSettled so a failed boot still assigns survivors for afterAll
cleanup); the env-specific fixtures never receive input, so their tests
only await their own readiness output. The slow-startup test now proves
frame ordering on the append-only output, which the old sequential
waits did not. File wall 10.2s -> ~4.9s, 3/3 repeat runs green.
* docs(ci): align gate and debounce descriptions with the removed admission job
* test(tooling): wait for readiness file content, not existence
writeFileSync creates the file before its bytes land, so the existence
poll raced the child's write on loaded runners and read an empty ready
file (observed in compact-small-4, run 29615028678). Poll for non-empty
content at both readiness sites.
* fix(mcp): bound doctor server checks
* fix(mcp): defer doctor concurrency helper loading
* fix(mcp): preserve doctor check coverage
---------
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(sessions): check file content size before Buffer allocation
Move the UTF-8 byte-length pre-check before Buffer.from() so oversized
session file content is rejected without first materialising the full
encoded payload. Buffer.byteLength is a native C++ call that counts
bytes without allocating.
Co-Authored-By: Claude <noreply@anthropic.com>
* test(gateway): prove oversized session write avoids buffer
---------
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Peter Steinberger <steipete@gmail.com>
* fix(infra): harden network client errors, session short ids, and extension load caching
* fix(infra): preserve injected dispatcher seams and drop the uncached extension loader
* fix(skills): filter NaN from Date.parse in history scan reviewed times
* fix(skills): filter NaN from Date.parse in history scan session cursors
runSkillHistoryScanCore built pending.sessionCursors by mapping
batch.sessions through Date.parse(session.updatedAt) without filtering
NaN. An invalid updatedAt timestamp would produce a cursor whose
updatedAtMs is NaN, breaking the resumed-candidate equality check
(candidate.updatedAtMs === cursor.updatedAtMs is always false for NaN)
and forcing every resume into the unreplayable-scan recovery branch.
Mirror the reviewedTimes NaN filter added in the previous commit by
dropping cursors whose Date.parse result is not finite. This keeps the
resumed-candidate matching predicate sound when an upstream feed ever
hands us a non-ISO timestamp.
* fix(skills): export toStoredState directly instead of __testing wrapper
The knip production deadcode scan flags __testing as an unused export
because test files are not entry points in that mode. Export toStoredState
as a named export instead of wrapping it in a __testing object, matching
the pattern of other test-accessible helpers in the codebase.
* fix(skills): remove toStoredState export to avoid knip unused-export flag
The knip production deadcode scan flags any export that is only
referenced by test files. Since toStoredState is not called from
production code, exporting it (whether as __testing or as a named
export) triggers the check-dependencies CI gate.
Remove the export and the toStoredState-specific regression tests.
The core fix (filtering NaN sessionCursors in resumePendingForBatchScan)
is still covered by the existing scan-cursor tests.
* fix(skills): remove unused SkillHistoryScanPromptSession type import
Cherry-picking the reviewedTimes NaN filter restored a type import
that was only needed by the original __testing-based unit tests.
Those tests were removed in a later commit, leaving the import
unused and tripping eslint(no-unused-vars).
* fix(skills): preserve history scan resume batches
---------
Co-authored-by: Patrick Erichsen <patrick.a.erichsen@gmail.com>
* [AI] fix(infra): archive plugin-state sidecar when canonical rows are strictly newer (#109832)
Conflict detection in migrateLegacyPluginStateSidecar compared rows for byte equality but treated every mismatch as a conflict. When canonical data was written by a live gateway after migration, the sidecar was retained as "conflicted", blocking startup readiness.
Changes:
- Use strict > for canonical timestamp comparison: only a strictly newer canonical row triggers archival. Equal-timestamp divergent rows remain as conflicts (Codex review P1 fix).
- Update error message: "have different values; sidecar data is newer" instead of misleading "already existed in shared state"
- 17 plugin-state migration tests covering all timestamp scenarios
Related to #109832
* fix(state): clarify plugin sidecar conflicts
---------
Co-authored-by: Josh Lehman <josh@martian.engineering>