Add crash-time ring flush and a diagnostics export UI (discussion #644, steps 4-5)

Stacked on the steps 1-3 branch (feature-diagnostic-logging, PR #646).
Kept as its own PR rather than folded into that one, matching the
discussion's own framing: step 4 is explicitly "the highest-risk piece
... lands last, behind its own switch."

## Step 4 -- crash-time ring flush (CrashHandler)

Installs a handler for SIGSEGV/SIGABRT/SIGBUS/SIGFPE/SIGILL (POSIX) /
SetUnhandledExceptionFilter (Windows) that flushes the in-memory ring to
a fixed crash_dump.log before the process dies.

This required reworking LogRing (step 3) to be genuinely lock-free, not
just mutex-protected: a signal handler that blocks on a lock the
crashing thread (or another thread) already holds turns a clean crash
into a hang -- no ring dump *and* no core dump, worse than doing
nothing. append() now claims a slot with a single atomic fetch-add;
dumpToFd() reads the preallocated entries directly and writes them with
write(2) only, looping on EINTR/short writes. Accepted tradeoff: at most
one entry can be read torn if a crash lands mid-append into that exact
slot -- documented in logring.h, and the alternative (a seqlock to
detect and retry) wasn't judged worth the complexity for that window.

Other invariants implemented per the discussion:
- sigaltstack with a static 64 KiB buffer, SA_ONSTACK -- a stack-
  overflow SIGSEGV has no usable stack for a handler without one.
- Nothing under the actual handler touches Qt, QString or the
  allocator: the dump path and a small header (version/git/OS/Qt) are
  precomputed into fixed char buffers by install(), which runs once at
  startup in normal context.
- Atomic test-and-set so only the first crash writes a dump; a second
  concurrent/nested fault goes straight to restore-and-re-raise.
- After writing, the handler restores SIG_DFL and re-raises (POSIX) /
  returns EXCEPTION_CONTINUE_SEARCH (Windows) so the OS's own crash
  path -- core dump, Windows Error Reporting -- still runs. A handler
  that "fixed" the crash by swallowing the signal would destroy exactly
  the post-mortem evidence this whole design exists to preserve.

Tested in this environment: POSIX/Linux only, all five signals. Sent
each directly to a running process and confirmed (a) crash_dump.log is
written with the correct header and ring contents, mode 0600, and (b)
the process still terminates via the signal with the kernel's own
"core dumped" flag set (exit code 128+signal, confirmed for all five).
The Windows path is implemented per the discussion's guidance but is
untested -- no Windows build available in this sandbox.

## Step 5 -- getting the data back out

- QETApp::checkCrashDump(), called from checkBackupFiles() only when
  there's no stale project file to recover this run (so the two
  prompts never both show, per the discussion), offers an unretrieved
  crash dump via DiagnosticsReportDialog and then deletes it regardless
  of the user's choice -- offered exactly once.
- A new "Aide > Enregistrer un rapport de diagnostic..." action
  (QETMainWindow) builds the same kind of report from the *current*
  session (QetLogger::buildDiagnosticsReport(): header + this session's
  log file) for a manual "attach this to a bug report" flow, not tied
  to a crash.
- Both go through QetLogger::redact() before ever reaching the user:
  the one redaction implemented is a literal replace of the home
  directory with "~", since an absolute path under it leaks the
  account name. The discussion's fancier "optionally redact project
  filenames too" isn't attempted -- reliably telling a project path
  apart from arbitrary log text is a much fuzzier problem than a
  literal prefix match.
- DiagnosticsReportDialog shows the full (already-redacted) content
  before saving, per the discussion: "the user is about to attach this
  to a public tracker."

Verified in a real GUI session (Xvfb): triggered a SIGSEGV, relaunched,
confirmed the crash-report dialog appears with the right header/content,
confirmed it does not reappear on a second relaunch, and confirmed the
manual "Save report" action produces a correctly-formatted report and
saves it to a chosen path.

Built clean, no new warnings.

## Build systems

Registered in both: cmake/qet_compilation_vars.cmake, and
qelectrotech.pro. The .pro needed explicit globs for the new
sources/logging/ui/ subfolder -- sources/logging/*.{h,cpp} was already
globbed, but unlike the other ui/ subfolders that one had no entry of
its own, so diagnosticsreportdialog.{h,cpp} would not have been built
under qmake.
This commit is contained in:
ispyisail
2026-08-03 16:08:17 +12:00
parent ea5117b148
commit 5dec36cb29
15 changed files with 758 additions and 74 deletions
+33 -25
View File
@@ -19,31 +19,33 @@
#define LOGRING_H
#include <QByteArray>
#include <QMutex>
#include <QVector>
#include <atomic>
#include <vector>
/**
@brief The LogRing class
Fixed-capacity, always-on in-memory ring of the most recent log
lines. Discussion #644 (step 3): the ring exists as forward-compatible
infrastructure for a future crash-flush (step 4, not implemented
here) as well as an on-demand "what just happened" snapshot, so its
entries are stored pre-formatted as plain bytes in storage
preallocated once at construction -- append() never allocates.
lines, preallocated once at construction -- append() never
allocates.
Entries are fixed-size slots rather than a byte-packed ring: with
kCapacityEntries * kEntryBytes chosen to land exactly on the 2 MiB
budget, this keeps wraparound trivial (whole-slot overwrite, so a
slot is always either fully the old entry or fully the new one --
no torn entries) at the cost of truncating any single line to
kEntryBytes, independently of the logger's own (larger) per-message
truncation.
Lock-free by construction, not just "thread-safe": step 4 (see
crashhandler.h) reads this ring from inside a POSIX signal handler,
where taking any lock is unsafe -- if the crashing thread happens to
be the one that already holds it (or any other thread does and never
gets scheduled again), the handler hangs forever, and you lose both
the ring dump *and* the core dump. So there is no mutex here at all:
append() claims a slot with a single atomic fetch-add, and
dumpToFd()/snapshot() read the preallocated entries directly.
Thread-safe via a plain QMutex. This is *not* the lock-free design
discussion #644 specifies for a signal-handler crash path (step 4)
-- no signal handler is installed by this code, so nothing calls
into the ring from inside a signal context.
Accepted tradeoff: if dumpToFd() runs while another thread is
mid-append into the exact slot being read (only possible in the
crash-handler case, and only for at most one slot), that one entry
may be read torn -- part old content, part new. Every other entry is
unaffected. This is deliberate: the alternative (a seqlock or similar
to detect and retry torn reads) adds real complexity for a window
that, per discussion #644, is not worth trading "the handler must
never block" against.
*/
class LogRing
{
@@ -54,24 +56,30 @@ class LogRing
LogRing();
/// Append one already-formatted, already-truncated log line.
/// Bytes beyond kEntryBytes - 1 are dropped with a truncation marker.
void append(const QByteArray &line);
/// Bytes beyond kEntryBytes are dropped with a truncation marker.
/// Never allocates, never blocks. Safe to call from any normal
/// (non-signal) thread concurrently.
void append(const QByteArray &line) noexcept;
/// Snapshot of the entries currently held, oldest first.
/// Snapshot of the entries currently held, oldest first. Normal
/// (non-signal) context only.
QVector<QByteArray> snapshot() const;
/// Async-signal-safe: writes every entry currently held to fd via
/// write(2) only -- no allocation, no Qt, no locks. May write a
/// torn entry under the rare race described above; never blocks.
void dumpToFd(int fd) const noexcept;
void clear();
private:
struct Entry {
char data[kEntryBytes] = {};
int length = 0;
char data[kEntryBytes];
std::atomic<int> length{0}; // 0 = not yet written this lap
};
mutable QMutex m_mutex;
std::vector<Entry> m_entries; // preallocated once, capacity fixed
int m_next_index = 0;
int m_count = 0;
std::atomic<quint64> m_write_cursor{0}; // monotonically increasing
};
#endif // LOGRING_H