mirror of
https://github.com/openclaw/openclaw.git
synced 2026-07-23 14:31:15 +00:00
Second calibration pass after the cli-runner reliability whale fix (#109772) and the stripe-wall correction (#109899): the compact group hints drifted below loaded-fleet reality, so FFD packed the heaviest groups into ~300s bins (checks-node-compact-large-2 at 242-314s shard wall, up to 389s job wall) while tail bins idled. Method: parsed [shard:*] begin/end timestamps from all 147 compact job logs across seven green CI runs whose head SHAs carry the whale-fix hints (runs 29605136624, 29605203485, 29605983019, 29606701461, 29611308972, 29611457693, 29611500865; 7 samples per group). Each hint is the per-group mean after dropping cache-warm/contention outliers outside [median/1.5, median*1.5]. Packing cap: with honest loaded walls a 220s cap no longer fits either pool (adds one job to each); 235s keeps the same job count (6 large + 15 small + dist) and flattens the ceiling. Median job setup overhead measured 60s (p90 87s), so a 235s bin stays near the 5-minute PR budget. Predicted bin shard walls (sum of measured group means): before: large 220/301/200/219/257/112, small heavies 251/235/218/217/215, max 301 after: large 235/232/229/228/228/156, small heavies 236/234/234/234/233, max 236 large-pool max/mean 1.38 -> 1.08 Biggest hint deltas (old -> new): startup-core 98->156, core-tooling-4 71->125, core-unit-fast-isolated 50->90, core-tooling-3 82->108, runtime-server 14->29, storage-state 55->70, media-ui 113->124, tui-pty 103->116, runner-cli-1 18->8, unit-src-security 108->95. Exclusive bins keep their 150s cap and 5-bin count: tui-pty (116) no longer shares a bin with tooling-isolated (166s combined measured); FFD now pairs tooling-2 with tooling-isolated at 144s. STRIPE_FILE_SECONDS_HINTS left unchanged: measured stripe walls sit in exclusive bins far below the packing ceiling (tooling stripes 94-125s), and refitting per-file hints would reshuffle stripe membership and invalidate the measured group means above.
52 KiB
52 KiB