The initial discussion
We were a few hours into a discussion about an electromechanical system, trying to improve its behavior, and by that point the whiteboard was full of curves, assumptions about timing, and competing theories about how the actuator was probably behaving under real conditions. The discussion was technically sound, and that was precisely the problem. We were getting very good at reasoning about something none of us could actually see.
At some point I asked the question that should have come up an hour earlier: do we actually know what the component is doing? Not what we commanded it to do, not what the model predicts it should be doing, not what we infer from how long we've been driving it. What is it actually doing? The answer, once everyone stopped to think about it honestly, was not really. And once that was on the table, a large part of the conversation changed shape.
Command is not state
This distinction sounds trivial until a system gets complicated enough that people start quietly treating the two as equivalent. You tell an actuator to move to a position, that's a command. Your software calculates how long the motor should run, that's a model. You assume the mechanical component reaches the expected position, that's an inference. None of those three things tell you where it actually ended up: friction changes, components wear, supply voltage varies, loads shift, tolerances stack up in ways nobody planned for, something gets partially blocked, a motor behaves differently once it's warm. Meanwhile the software can be completely convinced everything went exactly according to plan, because as far as it knows, it did.
This is one of the fundamental tensions of embedded systems, and I don't think it goes away no matter how good the tooling gets. Software lives in a world of discrete states and deterministic instructions. Hardware doesn't, and it never will.
Models are useful, until they quietly become a substitute for reality
There's nothing wrong with inference on its own, every engineering system leans on some version of it. We estimate temperature between samples, calculate position from encoder counts, approximate battery state of charge, derive flow from pressure differences, because measuring everything directly would be expensive, impractical, or simply unnecessary.
The problem starts when we forget that the model is still a model. In my experience the usual response to that kind of uncertainty is to make the model more sophisticated rather than to question whether a model is even the right tool for the job. If the actuator doesn't behave consistently, the instinct is to improve the timing curves. If those aren't accurate enough, add compensation for operating conditions. Then another correction factor, then calibration tables, then separate curves for different product variants, until the software contains a genuinely impressive representation of how we believe the physical system behaves. And the original question is still sitting there, unanswered: did it actually happen?
Sometimes the best improvement to an algorithm isn't a better algorithm. It's one more measurement.
What a feedback loop changes
There were several ways to close the loop, and they weren't interchangeable even though they initially sounded similar. We could observe something about the motor itself, measure the position of the actuator directly, or measure the physical result the actuator was supposed to produce. Monitoring electrical behavior might tell you whether the motor is moving or has hit a mechanical limit. A position sensor tells you exactly where the actuator sits. A flow sensor doesn't care about actuator position at all, it tells you whether the system is producing the result you actually care about, and that distinction turned out to matter more than anyone in the room had given it credit for.
It's common to measure the internal state because it's the easiest thing to measure, even when the real requirement is further downstream. If the objective is controlling flow, knowing valve position is useful, knowing the actual flow is better, and the closer the measurement sits to the outcome you care about, the fewer assumptions are left standing between what you observe and what's true.
Every unmeasured step is an assumption you're carrying
Picture a simple chain: command, then motor movement, then mechanical position, then physical output. If the software only knows the command, everything after it is assumed. Observe the motor and one assumption disappears. Know the mechanical position and another one goes. Measure the actual output and you've closed the loop around the thing that ultimately matters, though that doesn't automatically mean the last option is the right architecture for every system.
What it does is change the question worth arguing about. Instead of asking how we make our estimate more accurate, the better question is which uncertainty we're trying to remove. That question often does more work in a design review than another round of modelling.
The hardest part wasn't the sensor
Looking back, the hardest part of that discussion was never finding a way to close the loop. There were several technically plausible options, and engineers are generally comfortable comparing that kind of trade-off. The harder part was questioning an approach that had already existed for years, in a different team, with people who had built real expertise around it.
When a team has spent years controlling a system without direct feedback, a lot accumulates around that assumption: curves, calibration methods, test procedures, and people who have become genuinely good at making that model work. Eventually it stops feeling like one possible architecture among several and starts feeling like the way the system works, full stop. Then someone asks why we're estimating something we could just measure, and that lands as a much bigger challenge than it was ever meant to be.
At that point you're not only questioning an algorithm, you're questioning decisions made before you arrived and expertise people have built around them, sometimes on products that have already shipped. Evidence that challenges a new design is often easier to accept than evidence that challenges an established way of thinking, and I don't think that's a flaw in the people involved so much as what happens whenever expertise and identity get tangled together over enough years.
The discussion can slide from is feedback technically useful to why should we change something that has worked until now, and those aren't the same question, even though in the room they can feel like it. Something can have worked reasonably well and still carry an architectural limitation, and a product can be successful in the field while still containing assumptions we'd choose differently if we were designing it today.
There's real value in an established design, a system that has survived years in the field carries evidence a new proposal doesn't have yet, and replacing a proven approach purely because a cleaner architecture exists on paper would be irresponsible. But past success doesn't prove that every assumption inside that design deserves to stay untouched forever. The burden of proof runs both ways: whoever's proposing the extra feedback should say what uncertainty it removes and why removing it is worth the cost, but the answer from the existing design can't just be that we've always done it this way. That's history, not an engineering argument.
Part of the resistance is also that a measurement is less accommodating than a model. When expected and actual behavior disagree, a model leaves room to debate which assumption was wrong, whether the test was representative, or whether one more calibration curve would fix it. A measurement doesn't leave much of that room: some explanations survive and others quietly don't. That's exactly why observability is useful, and exactly why introducing it can expose weaknesses nobody could previously test directly. Good engineering culture has to be comfortable with that trade, because the point was never to prove one team right and another wrong. It's to make being right less dependent on who tells the most convincing story in the room. A measurement doesn't make everyone agree, but it gives everyone something harder to argue around.
Measurement isn't free
There's an obvious trap here, and it isn't that every system should just have more sensors. Sensors cost money, draw power, need connectors, PCB space, interfaces, calibration, diagnostics, manufacturing processes and their own software, carry their own tolerances, and can fail in ways that create new problems instead of solving old ones. Every sensor added to a product is one more thing that has to be tested, supported and maintained for the life of that product.
Sometimes an open-loop system really is the right design. Sometimes a simple model is accurate enough and adding observability would be solving a problem nobody has. Sometimes the economic cost of knowing something is genuinely higher than the cost of occasionally being wrong, and that trade-off gets sharper the moment you're talking about a product manufactured at scale rather than a bench prototype. Adding a €20 component to a lab prototype can look insignificant. Adding it to tens of thousands of units is a real decision.
But engineering effort isn't free either. Months spent compensating in software for something that could have been observed directly cost real money, and so does field troubleshooting on a system that can't tell you what actually happened. So the comparison was never sensor cost versus zero cost. It's cost of observation versus cost of uncertainty, and those deserve to be calculated explicitly rather than assumed.
Software can't recover information the system never captured
This is the part I think is easiest to underestimate, especially now. Software is extraordinarily good at processing information: filtering noisy signals, combining multiple measurements, spotting patterns, compensating for known behavior, and making increasingly sophisticated decisions on top of all of it. What it can't do is reconstruct information that was never present at the boundary in the first place. If the system has no observable indication of the physical state, the software can only ever estimate it, and a better algorithm produces a better estimate, not a measurement.
That distinction matters more now that increasingly sophisticated algorithms are being pushed onto embedded devices, because there's a real temptation to assume enough computation can compensate for limited observability. Sometimes it can. But intelligence doesn't remove physics, and if two different physical states produce exactly the same information at the software boundary, no algorithm is going to reliably tell them apart. At that point the limitation isn't computational. It's architectural.
It's the same mistake everywhere, not just in embedded systems
A software service with no useful logs has the same problem. A development process with no meaningful test evidence has the same problem. A business dashboard built from weak indicators has the same problem, and a team trying to speed up a process without measuring where the delays occur has it too. In every case, people compensate for missing information with theories, and sometimes the theories are genuinely excellent. But there's a point where another meeting, another spreadsheet, or another model provides less value than one well-chosen observation, and most engineers know that point when they hit it, they just don't always act on it.
That's probably why my first instinct, when something is uncertain, has stopped being how do we reason about this better and become what could we observe that would make the argument unnecessary. Not every disagreement can be solved that way. A surprising number of them can.
What I actually think the architect's job is
It's tempting to describe architecture as choosing technologies, drawing boundaries, defining interfaces, deciding how the software should be structured. Those things matter, but a large part of the job, maybe the largest part, is deciding what the system needs to know: which states need to be observable, which assumptions are acceptable, which ones are dangerous, where the loop gets closed, what information crosses the boundary between hardware and software, and where inference is genuinely good enough.
Those decisions shape everything that follows, and they often determine whether the software stays simple or ends up spending the rest of the product's life compensating for information it never had. I don't think the best architecture is the one with the most sensors, telemetry or diagnostics. It's the one with enough information to make the decisions that matter, no more than necessary, but no less either.
Where this leaves me
Asking whether the system should provide feedback didn't make the original discussion easy. There were still real trade-offs, different sensing approaches carried different costs, some required hardware changes, others touched the mechanical design, some were technically elegant and economically difficult to justify. But it became a better problem to be arguing about. We weren't discussing how to make our assumptions more sophisticated anymore, we were discussing which assumptions we wanted to eliminate altogether.
Because sometimes the right question isn't how do we predict this more accurately. It's why are we predicting it at all, when we could just go and find out.
0 comments:
Post a Comment