Inventory and status of zare systems at CoF

From dtype.org

Prepared 4 August 2026. All figures measured directly from the two servers that morning, not carried over from earlier notes. This page is unlisted and not linked from anywhere.

The short version

CoF's video archive lives on two servers, zare4 and zare5, holding six storage volumes totalling about 166 TB of usable space. All six are online and serving right now. Nothing has been lost.

Three things are worth CoF knowing, and they are not the three you would guess:

  1. The oldest drives are the healthiest. Eleven of zare4's twenty-four drives have been spinning continuously for 13.5 years. Every one of them reports zero reallocated sectors — the drive's own count of bad spots it has had to work around. That is genuinely remarkable for hardware of this age.
  2. The problem drives are the newer ones. Every drive failure and every meaningful sign of wear in the past two weeks has been on zare5's newest and largest array — the 10 TB drives bought in 2017. One died outright on 1 August and was replaced this morning; another is visibly wearing now.
  3. All three of the largest volumes are 100% full. This is the constraint that shapes everything else, and it is the one that eventually needs money rather than maintenance.

Nothing here needs an urgent decision from CoF. A drive replacement was completed this morning, a replacement spare is already ordered, and the next likely swap is identified in advance.

But do not read any of the above as "these machines are fine". They are in better shape than hardware this old has any right to be — that is not the same as being in good shape. These are 2012- and 2013-era machines running well past any reasonable service life. There is a timer on them and it is already past due. What we are doing is a maintenance game that has an end: the end is not in question, only the date is, and we cannot predict it. Replacing drives keeps the archive alive in the meantime; it does not make the machines younger.

What the systems are

Two servers in the rack, each with twenty-four hard drives plus a small boot drive. The drives are grouped into three volumes per machine, six in total. Each volume is RAID6, which means any two drives in a group of eight can fail without losing a single file. That is the safety margin everything below is measured against.

Server Volume Share name Built Drives Usable Used Status
zare4 /opt/fs1 zshare4a Aug 2012 8 × 3 TB 17 TB 100% Healthy, 8/8 drives
zare4 /opt/fs2 zshare4b Aug 2012 8 × 3 TB 17 TB 95% Healthy, 8/8 drives
zare4 /opt/fs3 zshare4c Apr 2015 8 × 6 TB 33 TB 100% Healthy, 8/8 drives
zare5 /opt/fs1 zshare5a Jun 2013 8 × 4 TB 22 TB 93% Healthy, 8/8 drives
zare5 /opt/fs2 zshare5b Jun 2013 8 × 4 TB 22 TB 85% Healthy, 8/8 drives
zare5 /opt/fs3 zshare5c Apr 2017 8 × 10–12 TB 55 TB 100% Rebuilding — 7/8 drives, see below

The volume names map to the shares you already use: a, b, c correspond to fs1, fs2, fs3. So a problem reported as "zshare5c" is the 55 TB volume on zare5 — the one currently rebuilding.

Current status, 4 August 2026

All six volumes are online and serving. Five are at full health.

zare5 /opt/fs3 (zshare5c) is rebuilding after this morning's drive replacement, and should be back to full redundancy by tomorrow. While it rebuilds, that volume tolerates one further drive failure instead of the usual two. The share stayed up throughout — the failure, the swap and the rebuild — and was readable and writable the whole time.

That is the whole story: a drive failed, a drive was replaced, the machine is copying data back onto it. It is routine, it is the process working, and it needs nothing from anyone.

Recent history

The past two weeks in detail. This is not because nothing happened before it — it is because detailed record-keeping only starts there, when we began systematically inventorying and measuring these machines. The years before that were not quiet, and the drive ages prove it.

These machines have been maintained for over a decade

The clearest evidence is in the inventory below. At least 22 of the 48 drives are still the ones originally installed; the rest have been replaced at some point over the past thirteen years. Roughly:

  • zare4 /opt/fs1 and /opt/fs2, built August 2012, still hold eleven of their original drives at 13.5 years each. Five others in those two volumes have been swapped in over the years, at 2.3, 2.6, 3.8, 4.8 and 7.1 years old.
  • zare5 /opt/fs1 and /opt/fs2, built June 2013, still hold six original drives at roughly 12.8 years each. The other ten have been replaced.
  • zare5 /opt/fs3, built April 2017, holds five of its original drives; the other three were replaced in the last two weeks.
  • zare4 /opt/fs3, built April 2015, appears to have been substantially turned over — no drive in it reads older than 6.9 years — but see the caveat below before treating that as a count.

So the picture is not "thirteen years of neglect followed by a bad fortnight". It is routine drive replacement, quietly, for over a decade — which is the reason a 2012 machine is still holding data at all. What changed recently is not the rate of failure but the visibility: we now measure every drive on both machines and write the numbers down, so replacements are planned from evidence rather than done reactively.

Worth saying plainly, because the table above is written in the passive voice and hides it: none of those replacements happened remotely. Every one of them needed somebody physically at the rack — powering a machine down, finding the right drive among twenty-four identical ones, swapping it, bringing the machine back up. That has been Jeannine's work, and this archive exists in 2026 because of it. The monitoring and the analysis can say which drive and when; they cannot pull it out of the shelf.

How many drives have been replaced? Less precisely than we first thought

An earlier version of this page carried a confident table here: 32 of 48 drives replaced, 67%, with two volumes at 100% turnover. That was wrong, and it is worth explaining why, because the reason is interesting and it affects how much weight to put on any drive-age figure.

The reasoning was: compare each drive's power-on-hours counter against the date its volume was built; anything younger than its volume must be a replacement. Sound in principle. The flaw is that on these older Seagate drives the power-on-hours counter is not reliable — the same defect already documented below for the 10 TB array turns out to affect six drives in zare5's 4 TB volumes too.

Those six report about 34,800 hours — call it four years — while simultaneously reporting nearly 113,000 hours of head-flying time (12.9 years), a figure that cannot exceed power-on hours, and a lifetime worst-case health index of 11 out of 100 that a four-year-old drive could not reach. Their power-cycle counts, 93 to 95, are the highest on the machine. Every independent signal says they are original 2013 drives, and only the hours counter disagrees. They were counted as replacements. They are not.

So the corrected picture, using head-flying time and power-cycle counts rather than the hours counter:

Volume Built Original drives still in place Replaced at some point
zare4 /opt/fs1 Aug 2012 5 3
zare4 /opt/fs2 Aug 2012 6 2
zare4 /opt/fs3 Apr 2015 see caveat see caveat
zare5 /opt/fs1 Jun 2013 3 5
zare5 /opt/fs2 Jun 2013 3 5
zare5 /opt/fs3 Apr 2017 5 3

At least 22 of the 48 drives are original to their volume — not 16 — and no volume has had all eight replaced. Around 18 to 26 drives have been changed over thirteen years, which is roughly one to two a year rather than the two to three claimed before.

⚠ The caveat on zare4 /opt/fs3. That volume is all Western Digital, and WD drives do not report head-flying time at all, so the cross-check that caught the error elsewhere cannot be run on them. Judged on the hours counter alone it looks fully turned over — but that is precisely the reasoning that just failed. Its power-cycle counts (19 to 61, against 100+ on drives known to be original) do suggest genuine replacements, so the volume has clearly seen work; we simply will not put a number on it.

The honest summary is that these machines have had steady drive replacement for over a decade, that more of the original hardware survives than we first credited, and that the exact count is not recoverable from the machines themselves. Only the drive currently in a bay is visible — a bay that has been through two drives looks identical to one that has been through one, and the earlier drive leaves no trace. Going forward this is a solved problem: every replacement from July 2026 onward is recorded with the date, the serial and the reason.

The past two weeks

Date Machine What happened
24 Jul zare5 A drive on the 55 TB volume was found dead. Replaced with a 10 TB spare; volume rebuilt and returned to full health.
29 Jul zare5 A second drive on the same volume, with over 16,000 reallocated sectors, was replaced before it could fail. Rebuilt onto a 12 TB spare.
30 Jul zare5 Rebuild completed clean. All three volumes healthy.
1 Aug zare5 A third drive on that volume died outright, at idle, with no warning. Volume dropped to 7 of 8 drives and stayed there — with no spare installed, it could not begin repairing itself.
1 Aug zare5 The monthly integrity scan was deliberately switched off on this machine while the volume was short a drive, to avoid adding a full day of heavy reading to the surviving drives. It is still off and will be switched back on once the rebuild finishes.
4 Aug both Both machines were powered down and back up (08:14 and 08:16) for the drive swap. Jeannine replaced the dead drive — a day ahead of the scheduled date — and the volume began rebuilding at 08:59.

Three drive replacements in twelve days is not normal, and it is fair to ask whether something is wrong beyond the drives themselves. Our assessment is that it is not: all three were on the same 2017-vintage array, all three are the same drive model, and the other forty-plus drives across both machines show nothing comparable. It reads as one bad batch reaching end of life together, which is a well-known pattern for drives bought at the same time from the same production run.

A note on how the failed drive was found

Worth recording because it worked and may be needed again. The dead drive was electronically unresponsive — it would not report its own serial number, so it could not be made to blink on command. There is also no working bay-number indicator on this chassis.

The method used instead: copy a large file off the affected volume and watch the shelf. Every healthy drive in that group lights up; the dead one is the only one in the group sitting dark. Identification by absence. The label serial is then checked before anything is pulled.

Jeannine got it right on the first attempt, working from nothing but a serial number and a description of which lights to watch. Pulling the wrong drive would have taken the volume to zero redundancy, so getting it right the first time genuinely mattered.

Complete drive inventory

Every drive in both machines, measured 4 August 2026.

How to read the columns:

  • Years — how long the drive has actually been spinning, converted from its own power-on-hours counter.
  • Realloc — sectors the drive has found bad and silently worked around. 0 is what you want. A low number that stays put is tolerable; a number that climbs is the warning sign.
  • Pending — sectors the drive suspects are bad but has not yet remapped. Usually cleared by the monthly integrity scan.

zare4 — 24 drives

Summary: zero reallocated sectors on all 24 drives. Eleven of them have been running for 13.5 years.

/opt/fs1 (zshare4a) — 8 × 3 TB, built August 2012

Drive Model Serial Years Realloc Pending Temp
sdb WD30EZRX WD-WMAWZ0265142 13.5 0 0 50 °C
sdc WD30EZRX WD-WMAWZ0324492 13.5 0 0 53 °C
sdd WD30EZRX WD-WMC4N0DDT38W 4.8 0 0 50 °C
sde WD30EZRX WD-WCAWZ2344412 3.8 0 0 51 °C
sdf WD30EZRX WD-WC81Y9051052 2.3 0 0 52 °C
sdg WD30EZRX WD-WMAWZ0361626 13.5 0 0 49 °C
sdh WD30EZRX WD-WMAWZ0408925 13.5 0 0 53 °C
sdi WD30EZRX WD-WMAWZ0342551 13.5 0 0 50 °C

/opt/fs2 (zshare4b) — 8 × 3 TB, built August 2012

Drive Model Serial Years Realloc Pending Temp
sdj WD30EZRX WD-WMAWZ0274013 13.5 0 0 49 °C
sdk WD30EZRX WD-WMAWZ0198152 13.5 0 0 52 °C
sdl WD30EZRX WD-WCAWZ2261213 13.5 0 0 51 °C
sdm WD30EZRX WD-WCAWZ2078646 2.6 0 1 53 °C
sdn WD30EZRX WD-WCAWZ2253727 13.5 0 8 54 °C
sdo WD30EZRX WD-WMC1T3071834 7.1 0 0 53 °C
sdp WD30EZRX WD-WMAWZ0324385 13.5 0 0 55 °C
sdq WD30EZRX WD-WCAWZ2261050 13.5 0 0 51 °C

The 9 pending sectors on sdn and sdm are the only blemish on this machine. They are small, they have not moved, and the monthly integrity scan is the normal mechanism for clearing them. Not a concern at this level; worth watching for growth.

/opt/fs3 (zshare4c) — 8 × 6 TB, built April 2015

Drive Model Serial Years Realloc Pending Temp
sdr WD60EFZX WD-C82DWNAK 2.0 0 0 46 °C
sds WD60EFPX WD-WX22D24DM29Z 0.3 0 0 41 °C
sdt WD60EFRX WD-WX11DB4H8KNY 2.5 0 0 48 °C
sdu WD60EFRX WD-WX11D58LZ72R 6.9 0 0 46 °C
sdv WD60EFRX WD-WX21D48FD47L 6.7 0 0 48 °C
sdw WD60EFAX WD-WX51D89HDFEA 4.7 0 0 49 °C
sdx WD60EFZX WD-C80T9S3G 3.9 0 0 51 °C
sdy WD60EFRX WD-WX11D36JRUAF 5.4 0 0 50 °C

zare5 — 24 drives

/opt/fs1 (zshare5a) — 8 × 4 TB, built June 2013

Drive Model Serial Years Realloc Pending Temp
sdb ST4000DM000 ZDH0515X 7.8 0 0 39 °C
sdc ST4000DM004 ZTT5XVJL 2.2 0 0 39 °C
sdd ST4000DM000 W30066MR ~12.9* 0 0 42 °C
sde ST4000VN008 ZDHAXA6M 3.2 0 0 39 °C
sdf ST4000DM000 W3005NVH ~12.8* 0 0 35 °C
sdg ST4000DM000 S300LKHH 1.5 0 0 37 °C
sdh ST4000DM004 ZFN5P3EV 0.6 0 0 35 °C
sdi ST4000DM000 W3006K96 ~12.9* 0 0 36 °C

\* Starred ages are derived from head-flying time, not the drive's power-on-hours counter, which has wrapped on these six drives and under-reports them by about nine years. See the replacement-count section above.

/opt/fs2 (zshare5b) — 8 × 4 TB, built June 2013

Drive Model Serial Years Realloc Pending Temp
sdj ST4000DM004 ZFN1MZA8 7.5 0 0 46 °C
sdk ST4000VN008 ZDHASAVA 4.3 152 0 44 °C
sdl ST4000DM000 W3005M66 ~12.8* 0 0 50 °C
sdm ST4000DM000 W3005QTA ~12.8* 0 0 50 °C
sdn ST4000DM004 ZFN5RQF0 0.2 0 0 40 °C
sdo WD40EZRZ WD-WCC4E6ZPE2R8 9.4 0 0 50 °C
sdp ST4000DM000 W30058V3 ~12.8* 0 0 43 °C
sdq ST4000DM004 ZFN1KWTX 6.7 0 0 44 °C

sdk's 152 reallocated sectors have been static for as long as we have been tracking. A fixed number is a drive that hit a bad patch once and worked around it — not a drive on its way out. Being watched, not scheduled for replacement.

/opt/fs3 (zshare5c) — 8 × 10–12 TB, built April 2017 — the array that has needed all the attention

Drive Model Size Serial Installed Realloc Temp Note
sdr ST10000VN0004 10 TB ZA21214K original (2017) 8 54 °C Static since first measured
sds ST10000VN0004 10 TB ZA210YA7 original (2017) 0 53 °C
sdt ST10000VN0004 10 TB ZA20CYVN 24 Jul 2026 0 56 °C Replacement
sdu ST10000VN0004 10 TB ZA20QMVH original (2017) 7,280 57 °C ⚠ The next one to replace
sdv ST10000VN0004 10 TB ZA20V0V8 original (2017) 0 52 °C
sdw ST12000VN0008 12 TB ZZ30LAEV 4 Aug 2026 0 57 °C Rebuilding now
sdx ST10000VN0004 10 TB ZA20TS1E original (2017) 0 53 °C
sdy ST12000VN0008 12 TB ZZ30LASR 29 Jul 2026 0 59 °C Replacement

Why this array has no "Years" column. The five surviving 2017 drives report their own age as roughly 860 hours — about five weeks — which is impossible for drives installed in 2017. The same drives simultaneously report nearly 3,900 hours of head-flying time, a figure that cannot exceed power-on hours, and a lifetime worst-case health index of 12 out of 100 that a five-week-old drive could never have reached. Their internal hour counter has plainly wrapped around or been reset, so we do not report it. The array was created on 30 April 2017, so the honest statement is that these drives are approximately nine years old. The three drives installed this year report their age correctly and are shown as installed dates instead.

We would rather say "this number is untrustworthy" than print a precise figure we know to be wrong.

Both machines also have a small 60 GB solid-state boot drive (zare4 and zare5, one each). These hold only the operating system, no video, and both report healthy.

Spares on hand

Size Qty Fits Note
20 TB 1 Earmarked for the Synology units, not for these two machines
14 TB 1 Same
12 TB 0 zare5 fs3 Used this morning on the failed drive. A replacement is ordered and expected on site this week.
8 TB 1 Does not match any array; usable only as an oversized substitute in a 6 TB or smaller volume
6 TB 2 zare4 fs3 Matched
4 TB 1 zare5 fs1 / fs2 Matched
3 TB 1 zare4 fs1 / fs2 Matched

Every volume has a matching spare on site except zare5 /opt/fs3 — the one that has needed all three replacements. Its replacement 12 TB is ordered and arriving this week, which restores cover.

And the "Fits" column is more forgiving than it looks. A larger drive can always stand in for a smaller one — the volume simply uses as much of it as it needs and ignores the rest. We have now done exactly that twice on zare5 /opt/fs3, putting 12 TB drives into 10 TB positions, and it works cleanly. The reverse is not true: a smaller drive cannot substitute for a larger one.

The practical effect is that the shelf is deeper than a strict size-for-size reading suggests. The 8 TB, for instance, covers any 6 TB or smaller position, and the 20 TB and 14 TB would physically work anywhere at all if it ever came to that — they are earmarked for the Synology units rather than reserved by any technical limit. So a spare being one size off is an inconvenience, not a gap in cover.

Diagnosis and what to expect

The old drives are healthier than they should be — which is not the same as safe

Eleven drives in zare4 have been powered on continuously since 2012 — about 118,000 hours each. Industry expectation for a consumer drive of that generation is somewhere between three and five years. These are at thirteen and a half, and every one of them reports zero reallocated sectors.

That deserves a caveat rather than a victory lap, and the caveat is the important half. Zero reallocated sectors means the drive has not yet had to work around a bad spot. It says nothing about bearings, motors or the mechanical wear that actually ends a drive of this age, and it gives no warning at all — the drive that died on 1 August had counters that looked fine until the moment it stopped answering. A drive at thirteen and a half years can fail suddenly, without warning, and the honest expectation is that these will begin to.

So the correct reading is narrow: there is no evidence any of them is failing right now, which is a statement about today and not a forecast. We are not replacing them pre-emptively because there is nothing to act on and eleven working drives is real money — not because we expect them to last. We are running on borrowed time and choosing, deliberately, to keep borrowing it. RAID6 is what makes that a reasonable choice rather than a reckless one: it lets us wait for evidence instead of guessing, and it means a surprise failure costs a rebuild rather than an archive.

The wear is on the newer drives, not the old ones

This is the counterintuitive part and the most useful thing in this report. Every failure and every meaningful sign of wear across both machines is on zare5's 2017-vintage 10 TB array:

  • Three drives from that array replaced in the past twelve days.
  • One surviving drive, sdu, at 7,280 reallocated sectors and climbing — up roughly 800 since the end of July.
  • Meanwhile all 24 drives in zare4, including the 13.5-year-old ones, are at zero.

Age is simply not the predictor here. Drive model and manufacturing batch are.

What to expect next

  • Tomorrow: the rebuild finishes and zare5 /opt/fs3 returns to full two-drive redundancy. The monthly integrity scan is switched back on at that point.
  • ⚠ Expect another drive replacement on that same volume soon — drive `sdu`. This is a firm expectation, not a worry. It is at 7,280 reallocated sectors and actively climbing — up roughly 800 in the past few days, having been flat before that. It is the only drive on either machine that is currently moving. We are deliberately not stacking that swap on top of the current rebuild, because doing two at once on the same volume narrows the safety margin for no good reason. Once the replacement 12 TB drive arrives this week and the current rebuild has finished, this is the next piece of work, and we will contact Jeannine to arrange it.
  • Beyond that, no specific prediction. The remaining 2017 drives could last years or fail next month; the honest answer is that we will see it in the numbers first, if we are watching, which brings us to the one real gap.

The one genuine gap in our monitoring

On zare5, the automatic drive-health monitoring does not cover the 10 TB array — the exact array that has had all three failures. The monitoring software's configuration lists the machine's first seventeen drives and stops short of the eight that make up that volume. This is a long-standing oversight, not a recent change.

The practical consequence: on that array we find out a drive has already died, never that one is about to. The drive that failed on 1 August had been quietly accumulating damage for months, and nothing flagged it. The two drives we caught early were caught by hand, during a manual check.

Separately, neither machine has ever run the drives' own built-in self-tests — there is no schedule configured for them, on either box, going back years.

Neither of these has caused a data loss, and RAID6 is doing its job. But they are the difference between planned replacements and surprise ones, and both are on our list to fix. No action is needed from CoF on this — it is our configuration to correct.

Capacity: the real constraint

zare4 /opt/fs1, zare4 /opt/fs3 and zare5 /opt/fs3 are all 100% full. The other three are at 85–95%.

This is not causing failures, but it does mean:

  • There is nowhere to stage or copy files during maintenance.
  • New recordings have nowhere to go on those volumes.
  • We cannot expand these volumes in place. The machines are thirteen years old; putting larger drives in them one at a time would take years to yield any extra space and would mean investing in hardware at the end of its life.

And this is where the timer meets the money. The capacity problem and the age problem have the same answer, which is why they are worth thinking about together rather than separately. Every drive bought for these machines from here is life support for hardware that is already past its due date — worth doing, because the archive is irreplaceable and the machines are still working, but it buys time rather than progress.

The sensible path is new storage rather than more drives for these machines — a further Synology unit with modern 20–28 TB drives, which is how the newer storage at the site is already built. That is a budgeting matter for CoF and there is no urgency attached to it. In the meantime these machines will be kept running, with drives replaced as they fail.

Summary for CoF

  • Your data is intact and all six volumes are serving. No files have been lost at any point.
  • One volume is rebuilding after this morning's drive replacement and should be back to full redundancy by tomorrow.
  • Both machines are well past their design life and are performing better than they have any right to — particularly zare4, where eleven drives have run 13.5 years with no errors at all.
  • ⚠ That is luck and maintenance, not durability, and it will not hold indefinitely. The timer on this hardware is already past due. We are playing a maintenance game with an end; we simply cannot say when it arrives. Nothing about the current good news changes that, and it is the reason the longer-term storage question is worth having on the table before it becomes urgent.
  • One more drive replacement is expected shortly on zare5's large volume. We will reach out to arrange it once the replacement drive arrives.
  • Spare drives are on the shelf for every volume, with a replacement 12 TB arriving this week. Sizes do not have to match exactly — a larger drive always works in a smaller position.
  • Nothing needs a decision from CoF today. The only open item is the longer-term storage question, which is a budgeting conversation whenever CoF wishes to have it.

Thanks in particular to Jeannine. Three drive swaps in twelve days, each one identified by serial number under awkward conditions, none of them wrong — on top of years of doing exactly this. Everything in this report that looks like good health is downstream of somebody being willing to go stand at the rack.