CS-06Product designVertiv iCOM-SThermal management
How Long Do I Have
A dashboard that showed you everything, and told you nothing. Designing severity into a data center thermal display.
A data center room has a temperature it is supposed to stay inside. When it drifts out, you have some amount of time before hardware starts failing, and it is not a lot of time. So the question the person watching the screen actually needs answered is not what the temperature is. It is how bad this is, and how long they have.
The existing display could not answer either one. It presented the data; that was the whole design. Everything laid out uniformly, in the same fashion, every number weighted the same as every other number. Severity was not something the interface told you. It was something you worked out yourself, from the numbers, in your head, while the clock ran.
You had to already be a trained technician to read it: to know which alarm indicator referred to what, how serious it was, what would resolve it, and how long that resolution would take before it stopped mattering. That is a lot to ask of an interface whose entire job is to tell you when something is wrong.
Asking the people who live in it
The release was scoped around two questions the business had already framed: how to bring red, yellow and green health into the product, and what a thermal feature pack should contain. Both had answers already circulating internally, and not one of them came from a user. So before designing anything I built a workshop program to answer them with the people who would use the result.
I designed and facilitated four workshops across three user groups: two external data center operator customers, and the company's own field service technicians. The Product Manager and Product Lead brought the client relationships and the business framing; the research design, the activity construction, and the facilitation were mine.
Four activities, each doing a specific job.
Affinity diagramming for task and feature discovery, built around blunt questions rather than polite ones:
- What are the top five tasks you need this to do to make your job easier?
- What is the most frustrating thing you face when dealing with our products?
- What takes too long, and how much time would you get back if it did not?
Teams rarely know what they are missing, so asking directly gets very little; asking about friction gets everything.
Build your own interface, where participants laid out the screen themselves. What they put at the top, and what they buried, is information you cannot get by asking someone to rank a list.
Dot voting across the four phases of the equipment lifecycle (design, deployment, commissioning, operations), so prioritization came from the room rather than from me.
A user intersection map, because the thing I most needed to understand was not what any one group wanted; it was where the groups overlapped. Service technicians, commissioning agents, and customer operations crews all touch the same system at different moments, and the general contractor driving the job schedule sits outside all three while setting the pace for everyone. Overall health and reporting sat in the intersection. That told me what the interface had to serve first.
Designing for duration, not just state
Working from the workshop output, I needed the technical parameters before I could design anything real. Product had the established relationships with engineering, so they went and got them, and we walked through all of it in meetings until I understood the system well enough to design against it. From there, rough wireframes, then several sessions turning those into interface designs.
Three questions kept surfacing, and they turned out to be the same question wearing different clothes.
- How do you show when an event occurred, on a timeline?
- How do you tell a non-alarm event apart from an alarm event?
- How do you show how long something stayed in an alarm state?
All three are about time. The old display was a snapshot, and a snapshot cannot tell you whether you are looking at a blip or at something that has been degrading for eighteen hours. So the design moved from reporting current state to reporting duration and trajectory: time outside thresholds rather than temperature, alarm duration rather than alarm count, an event timeline that distinguishes a set point change from a component staging from an alarm state change.
Solving it once instead of eleven times
The early concepts were full-screen mockups, and about a third of the way in I stopped making them.
Two reasons. The practical one was pace: full screens are slow, and mocking every screen in a mature application to prove a handful of ideas is expensive. The design reason mattered more. The same problems were recurring across the application, which meant the useful unit of work was not a screen. It was a pattern.
So I shifted to designing specific zones of the interface as modular tiles, each with a simple state and an expanded state: the simple tile carrying the one number that tells you whether to care, the expanded tile carrying the evidence behind it. Configuration verification, set point range compliance, active alarms, alarm duration by unit; different content, same structure. A pattern solved once and applied wherever the situation recurred, which scaled the effort toward whatever was most impactful and gave the product a consistent grammar instead of eleven bespoke solutions.
Product liked the concept of establishing patterns that repeat. I would end up making a formal version of this argument at a later organization, as a governance system rather than a set of tiles, but the instinct started here.
Closing the loop
Concepts went back to people who had been in the workshops for review and feedback, which is the part I care about most. The questions they raised in a workshop are only worth asking if the answers get shown back to them; otherwise the workshop is theater.
What I did not run on this project was formal usability testing against tasks. The validation was structured design review with the users who had named the problems, not measured task performance. Worth being precise about, because the two get conflated constantly.
The outcome
Design was handed off to engineering to build; that was how the organization worked. Requirements reached me through Product, concepts went back to engineering for feasibility review, and then the design went off to be built. What I did not have was a direct line to the people doing the building. I rarely heard from developers unless they had a question, and every technical answer I needed arrived secondhand. Getting design and engineering into the same conversation without Product standing in the middle as translator was something I was working on when my role there ended in October 2022.
So I do not know what shipped. I know what the users said they needed, what I designed in response, and what the design was trying to buy them: the ability to look at a screen and know, without already being an expert, how bad it is and how long they have.
What I did
Designed the research program and facilitated four workshops across two external customers and internal field service, running affinity diagramming, a build-your-own-interface exercise, dot voting and user intersection mapping. Designed the wireframes and the high fidelity concepts, and made the call to shift from full-screen mockups to a reusable simple-and-expanded tile pattern. Took the concepts back to workshop participants for review. Worked with Product to get the thermal parameters I needed out of engineering, and learned the system well enough to design against it.
Prototypes and workshop material from 2022, shown as they were. Customer names and site identifiers removed.