SolarWinds should provide a native historical view that explains why a node changed its status at a specific point in time.
Current situation:
A node may temporarily change from Up to Warning, Critical, or Down and return to Up before an administrator investigates it. The Events resource records the timestamp and the status transition, but it does not retain or display the reason that caused the historical status.
The current cause may be visible in a tooltip or status-details view while the node is still affected. Once the node has recovered, administrators must manually correlate multiple metrics and child objects around the timestamp.
SolarWinds Technical Support confirmed in case #02171703 that there is currently no out-of-the-box widget, report, resource, or Node Details view that reconstructs the historical root cause after recovery.
Requested functionality:
Provide a native widget, report, or historical status-details view that records and displays:
- Exact timestamp of the status transition
- Previous and new node status
- Object, child entity, or metric that caused the status
- Threshold or condition that was breached
- Relevant measured value and configured threshold
- Whether the cause was packet loss, response time, CPU, memory, interface, volume, application component, hardware sensor, or another contributor
- Timestamp and condition of the recovery
- Links to the affected object and historical metric chart
The information should be available independently of alerting. A status transition may occur even when no alert action or alert definition exists.
Example:
01:24 – Node changed from Up to Warning
Cause: Memory utilization
Measured value: 94%
Warning threshold: 90%
Affected object: Node memory
01:29 – Node returned to Up
Measured value: 82%
This would significantly reduce troubleshooting time and provide administrators with a clear audit trail for temporary or overnight status changes.