MobAIThread and MobRespawnThread both ran `while (true)` around a sweep of every
zone with no delay at all, so each of them pinned a full core for the lifetime of
the world server. Measured in the container: 195% total CPU with
tid 229 aiThread 106% of one core
tid 231 respawnThread 106% of one core
and nothing else above 1%. The 100 ms check in the respawn loop only throttles how
often a mob may respawn, it never paused the loop.
Both sweeps now run at a fixed rate (AI_TICK_INTERVAL_MS / RESPAWN_TICK_INTERVAL_MS,
100 ms by default) by sleeping the remainder of each tick. MobAI.DetermineAction
already gates individual actions on timestamps (lastAttackTime, nextCastTime,
patrol delays), so pacing the sweep removes the busy spin without changing the AI
behaviour; raise the constants to trade reaction time for CPU.
LoginServer.exec() -> checkServerHealth() -> isPortInUse() shelled out to
`lsof -i tcp:<port>` through /bin/bash and read the child's stdout until EOF.
There is no timeout and no waitFor(), so a child that blocks is fatal: in the
container lsof was observed stuck in uninterruptible I/O (D state) while it was
scanning the busy world server. The login server then never returned from its
startup path, the accept loop never started, and clients could open a TCP
connection to 6000 and wait forever for the 100 byte DH handshake, surfacing as
"Failed to open a server connection" after the client's own timeout. A container
restart was the only way out.
Replace the external probe with a bind probe: it cannot hang, needs no external
tool, and SO_REUSEADDR still does not allow stealing a port that is actively
listening, so a successful bind keeps meaning "free".
The stock scripts stop the game with SIGKILL (mbkill.sh runs kill -9) and
'./reboot' only calls mbrestart.sh, so anything still parked in a character's
deferred database job is thrown away: experience is written five minutes after
it is earned and stats/skills thirty seconds after they change
(engine.jobs.DatabaseUpdateJob). PlayerCharacter.updateDatabase() is an empty
stub, so there was no write-out path at all.
Added engine.gameManager.GracefulShutdown:
flushAllPlayers() - writes skills/powers, stat modifiers and experience
for every online character
announce(String) - server wide flash message
registerShutdownHook() - flush on an orderly SIGTERM (SIGKILL cannot be
intercepted by design)
gracefulStop(int) - countdown, flush, then stop the login and world
servers
Added the './shutdown [seconds]' dev command (engine.devcmd.cmds.ShutdownCmd,
default 60s, ADMIN only like every other dev command), registered together with
the shutdown hook in WorldServer.main and LoginServer.main.
Verified on a MagicBox container: ant builds clean, the boot log shows
"shutdown hook registered for WorldServer"/"LoginServer", and kill -TERM on the
world server logs "WorldServer is terminating, flushing player data" ->
"flush finished, saved=0 failed=0" -> clean exit.