Municipal water and wastewater

It read the plant's control rules off the readings, to two centimetres

One small SCADA, an RTU reaching remote stations over radio, and an integrator who set the naming up and has since retired.

INTAKEraw waterTREATMENTdosing, filtersRESERVOIRlevel, turbidityNETWORK PUMPSpressure zonesREMOTE STATIONSRTU over radioone small SCADA · an RTU point list · one integrator, retired

The least documented site on this page, and the one the regulations now ask the hardest questions of.

What actually goes wrong

A municipal water works has the smallest engineering team and the oldest naming conventions of anything on this page. The person who set the tag names up was an integrator on a project fifteen years ago. They are not coming back, and what they knew was never written down anywhere that survived.

Two things made this urgent rather than merely untidy. NIS2 puts reporting obligations on operators who have no one to write the report, and the attack on the dam at Bremanger in 2025 moved OT resilience from a slide to a board question. Both ask the same first question: what have you actually got, and does it still work?

That question is answerable from tags, values and whatever paperwork survived, which is exactly what the compiler eats. This page used to say we had not run on water and were not going to pretend otherwise. We have now run on water.

The network is C-Town, published with the BATADAL competition under CC BY 4.0: a benchmark distribution network based on a real medium sized town, with a year of hourly SCADA readings simulated from its hydraulic model, and the model published beside it. It is not a working utility's data, and we have not yet run on one. 43 tags, 7 tanks, 11 pumps. The model file is the paperwork and the compiler never opens it. It is used to mark the answers, the way the HAI manual is on the rig page.

The ordinary result first. With no water connector at all the compiler places nothing and refuses all 43 tags, which is the same answer it gives any sector it has never seen and is the honest shape of this product. With a connector, which is a 78 line text file and not a release, it places 39 of 43 and gets every machine right. Name matching also gets every machine right, because this network writes its element names into its tags, and on the quantity it beats us 100 to 93. That criterion failed and it is reported as failed.

The reason it failed is the more interesting half. The three tags we lost are F_PU3, F_PU5 and F_PU9, and all three are flat zero for every one of the 4,177 hours. They are standby pumps that never start. The compiler reads values as well as names, and a series that is nothing but zero looks exactly like an off signal, so it called them statuses. Name matching is right about them only because it never looks: it reads the F and says flow, and on a dead channel that guess happens to land.

Then the part with no competitor. Which pump fills which tank is not written in any tag name. F_PU2 and L_T1 share no character that means anything, so there is no regular expression, no dictionary and no language model reading names that can connect them. It is in the numbers: a tank that rises while a pump runs, for a year. Asked that question with the model file shut, the compiler got six of six, and the control logic written by the people who built the network agrees with every one.

One tank it should have refused and did not. T6 has no monitored pump, and the compiler named PU10 at a correlation of 0.53. That is not noise: PU10 discharges five pipes from T6 and ten from T7, the tank its controls actually name, so T6 sits on the main PU10 pumps into. The answer describes the hydraulics correctly and contradicts the control rule, and the criterion said refuse. It is written down as a failure because that is what it is.

That failure is now fixed, and fixing it produced something better than the thing it fixed. Correlation was answering the wrong question: it finds the district, and a tank a pump fills looks identical from the outside to a tank that merely sits downstream of one. But a controlled tank does something a downstream tank cannot. It commands the pump, and the command is a threshold, so at every single switch-on the commanding tank's level is at nearly the same value while every other tank's level is wherever it happens to be. The signature is not correlation, it is tightness.

Tightness recovers the number as well as the pairing. Not that PU2 is associated with T1, but that PU2 starts when T1 falls past 1.0 metres and stops at 4.5. On a file it had never seen, 5 of 5 pumps that cycle enough to test were matched to the tank that commands them, T6 was refused, and the 10 recovered thresholds came back within a median of 2 centimetres of the numbers in the network's own model. 3 of the 10 landed on the stated figure to the centimetre.

Getting there took three preregistered runs and one of them failed. The first estimator was biased, always toward the middle of the tank, because the moment a pump starts the tank stops falling and begins to rise, so the reading after the switch sits above where the tank was heading. The fix for that bias was noisier, and the same number was being used both to pick the tank and to report the threshold, so the accurate version started refusing pumps the biased one had matched. Choosing wants a steady number and reporting wants an unbiased one, and they are not the same number. Each run is written down, including the one that went backwards.

The run, on a network nobody here designed

C-Town, published with the BATADAL competition under CC BY 4.0. A year of hourly simulated SCADA on the left of the scoring line, the network's own hydraulic model on the right of it. The compiler never opens the model; it is what marks the answers.

r 0.63r 0.82r 0.90r 0.56r 0.85r 0.53r 0.64PU1PU2PU3PU4PU5PU6PU7PU8PU9PU10PU11V1V45V47V2T3T1T7T6T5T2T4

C-Town as its own model draws it: 429 mains, 7 tanks, 15 pumps and valves, at the coordinates the benchmark gives. The curved lines are the compiler's answers to which flow fills which tank, with the correlation that produced each one. Six are right. The amber one is T6, which should have been refused.

Placing the tags

setupplacedquantitymachine
nothing: no connector, no model0 of 4328%0%
no connector, but the element list11 of 4328%67%
the water connector, tags only42 of 43100%100%
the water connector and the list42 of 43100%100%
name matching, the baseline—100%100%

Scored against the element the network's model names, never against the tag text. With no water connector the compiler places nothing, which is the honest shape of a sector-blind core. The connector is 78 lines of text and no compiler code changed.

Which pump fills which tank

tankthe compiler saidthe controls say correlationverdict
T1PU2PU1 or PU20.63agrees
T2V2V20.82agrees
T3PU4PU4 or PU50.90agrees
T4PU7PU6 or PU70.56agrees
T5PU8PU80.85agrees
T6PU10nothing monitored0.53missed
T7PU10PU10 or PU110.64agrees

From the values alone, with the model file shut. 6 of 6 tanks correct on held-out data. Name matching scores zero here and cannot score otherwise, because the relationship between a pump and the tank it fills is not written in either name. T6 has no monitored pump and should have been refused; it was not, and that criterion is recorded as failed.

How far we have got

the label on this story

measured

C-Town, published with the BATADAL competition under CC BY 4.0. With no water connector it places 0 of 43 tags; the connector is a 78 line file and no compiler code changed. Four preregistered runs, each on data the one before had not seen, and the two that failed are on this page.

read next

The numbers, including the ones that missed

Four preregistered runs, P17 and P22 to P24, each on data the one before had not seen. The reports are in auge_plant/, including the two criteria that missed and the run that went backwards. Reproduce the lot with three curl commands and two modules.