Skip to content
THE OBSERVINGFIELDBOOK

New field records

Recently added to this route

  • An Observing Log on a MacA season of visual notes needs a table, a client that opens it without ceremony, and a job that copies it while the observer sleeps.
  • PC Stability and Diagnostics: A Field MethodMemory tests, SMART attributes, sensor readings and burn-in each answer a narrow question. A method for reading PC hardware diagnostics without overclaiming.
  • Teaching Windows Upkeep as a SkillA field guide to teaching Windows maintenance, security, and hardware habits as practical skills for students and families who share a home computer.

FIELD GUIDE

PC Stability and Diagnostics: A Field Method

Memory tests, SMART attributes, sensor readings and burn-in each answer a narrow question. A method for reading PC hardware diagnostics without overclaiming.

01

Start with the useful answer

A PC that crashes under load is not diagnosed by a single test. Each tool answers a narrow question: memory tests check whether a specific pattern survives a write and read cycle, SMART attributes report counters the drive firmware chose to expose, sensor readings show what a chip reports at one moment. None of them proves a machine is stable. They only fail to prove it is broken. A method that treats each result as partial evidence, and records what was tested and for how long, is more useful than any single pass or fail.

02

What a memory test actually proves

Memory diagnostics write patterns to RAM and read them back. MemTest86 and its variants run several such patterns, from simple walking bits to moving inversions. A clean pass means the tested addresses returned the expected values under the tested patterns at the tested timings. It does not mean the modules are good.

Three limits matter. First, coverage: a short run may touch only part of the address space, and errors often sit in a narrow band that a longer run would reach. Second, conditions: a module that passes at stock voltage and 2133 MT/s may fail at 3600 MT/s with tightened timings, because the test did not reproduce the load. Third, temperature: DRAM errors rise with heat, and a test run in a cold case says little about behaviour after an hour of gaming.

A practical rule: run at least four full passes, then repeat with the case closed and the GPU loaded if the fault appears under use. Record the exact speed, timings and voltage. A pass at different settings is a different result. The same discipline applies to the wider diagnostic workflow described at the burn-in desk, where each test is tied to the condition it was run under.

A open desktop PC case on a workbench at night
Memory tests, SMART attributes, sensor readings and burn-in each answer a narrow question.
03

How to read SMART attributes without overreading

SMART is a set of counters defined by the drive manufacturer and reported through the firmware. Tools such as smartctl or CrystalDiskInfo display raw values and, sometimes, a normalised score. The raw value is the one to read.

Some attributes are well understood. Reallocated sector count (ID 5) counts sectors remapped to spare area. Current pending sector count (ID 197) counts sectors that failed a read and await remapping. Offline uncorrectable (ID 198) counts sectors that failed offline scanning. A non-zero value in any of these is a signal, not a verdict. A drive with two reallocated sectors and a stable count over months behaves differently from one gaining sectors weekly.

Other attributes are vendor specific. A raw value of 100 in one field may mean nothing in another brand. Normalised scores are scaled by firmware and can hide a rising raw count. The useful practice is to log raw values at intervals and watch the slope. A single snapshot is a photograph; a series is a trend.

SMART also does not cover everything. It says nothing about the controller's cache behaviour, the quality of the SATA or NVMe link, or the filesystem above it. A drive can report clean SMART and still drop a link under load.

04

What do sensor readings tell you?

Sensors report what a chip measures at a point in time. Motherboard tools read voltage rails, fan speeds and temperatures through a Super I/O chip. CPU packages report per-core temperatures through digital thermal sensors. GPUs report their own set.

Each reading has a tolerance. Voltage rails are typically within plus or minus five percent of nominal, and a 12 V rail at 11.6 V is inside that band. A reading of 11.4 V is not, but the meter itself may be off by one or two percent. Software sensors are also sampled at intervals, so a spike between samples is invisible.

Temperature readings depend on where the sensor sits. A CPU package temperature reflects the hottest core, not the cooler base. A VRM sensor may sit near one phase. Comparing a reading to a review of a different board is not a comparison.

The useful approach is to record idle and load values, note the delta, and check whether the delta grows over time. A rising delta at the same ambient and fan curve points to dust, dried thermal paste or a failing fan. A stable delta is a baseline, not a guarantee.

05

What does a stress test prove?

A stress test applies a sustained load and checks whether the system completes it without error, throttle or crash. Prime95, OCCT, y-cruncher and similar tools differ in the instruction mix they generate. Some hammer AVX units, some float, some integer. A system can pass one and fail another.

Passing a one-hour test proves the machine survived that hour under that workload. It does not prove stability for a different workload, a longer session or a warmer room. Failing a test proves the configuration is not stable under that load, which is a stronger result.

Burn-in duration is a choice, not a standard. Many overclockers use a short pass to screen obvious faults, then a longer run of several hours to catch thermal drift. The longer the run, the more confidence, but confidence never reaches certainty. A fault that appears once a week will not show in a four-hour test.

Reading an instability matters as much as detecting it. A worker thread that stops with a rounding error points to core or memory. A system that reboots without a log points to power or a hard fault. A test that throttles but completes points to cooling. Each symptom narrows the search.

06

Why each result has a limit

Diagnostics are bounded by three things: coverage, condition and duration. Coverage is how much of the system the test exercises. Condition is the voltage, clock, temperature and workload present during the run. Duration is how long the run lasts.

A result is only valid inside the box those three define. A memory pass at stock settings says nothing about an overclocked profile. A SMART snapshot says nothing about next month. A temperature reading says nothing about a spike between samples. A stress test says nothing about the workload it did not run.

The practical consequence is that a diagnosis is a record, not a verdict. Write down the settings, the duration, the ambient temperature and the exact tool version. When a fault returns, the record tells you whether the condition changed or the hardware did.

07

Building a repeatable method

Start with the symptom. A crash under gaming, a boot failure, a random reboot and a slow file copy point to different subsystems. Test the subsystem the symptom implicates first.

Run memory tests at the settings in use, not at stock, if the fault appears under load. Log SMART raw values before and after. Record idle and load temperatures and voltages. Run a stress test long enough to cross the thermal steady state, which often takes twenty to thirty minutes. Repeat the failing condition rather than a convenient one.

Keep a live USB with the tools on hand so a suspect machine can be tested without its own operating system. The method matters more than the toolkit. A short, consistent record of what was tested and what happened is worth more than a long list of passes with no conditions attached. Diagnostics answer narrow questions, and the same discipline applies away from the bench. A SMART attribute, a memory error count, a sensor reading at load: each one is a single measurement with a known tolerance, and the method is to record it, compare it against a baseline, and only then decide. Collectors use the same habit when they inspect a finished model, checking seams, decals and paint under steady light before any money changes hands. The model replica shops guide sets out those checks, along with storage and display practice for aircraft and ship models.

Source note. Practical context is checked against the primary and specialist records in our source register, then bounded by the methods recorded on the methodology page.

Keep your bearings

TelescopesBuild a usable telescope system from mount to finder and eyepiece.

Used GearInspect a used telescope, mount, or binocular without relying on a listing description.

What Is Astronomy?A practical introduction to astronomy as a science and a field practice.

Astronomy for BeginnersStart with your own sky, then add tools only when they make the next observation easier.