Codec Debug Session 2: Tell Us What You Really Mean

The appetizingly audible cracking open of the aluminum can was followed by a satiating, ice cold, citrus-y burst of carbonation as Carl sipped on his “something fizzy”. Something about carbonated water satisfies thirst in a different way than still water does. He referenced the list of topics that he had initially identified for remediating the decompression bug in the Techno-Wisdom Codec. He marked through the first two topics he had previously addressed leaving two more vying for his attention.

  1. The use of techno-wisdom to sell a utopic vision into an organization.
  2. The concerns surrounding the downstream incentives of those that would be tasked with executing such a vision.
  3. The ease with which the techno-wisdom can be used to express very real shortcomings of existing data and data systems.
  4. The unaddressed impracticalities of migrating from a legacy state to the acclaimed ideal state of a fully centralized data system.

Carl caught himself starting to feel cocky as he observed the significant progress he had made on the bug fix. However, he consciously paused, recalling the cautionary message on the ticket sitting on the surface in front of him, “….proceed with caution”. Glancing at the clock and trying not to underestimate the challenge ahead, he committed to addressing only one more topic before he left for the day.

What Do People Really Mean When
They Ask for “One Place for All Our Data”?

Carl placed his carbonated beverage a safe distance from his terminal equipment to prevent any inadvertent spills while working. Resting his fingers on the keyboard, he slowed his breathing—this next topic would require nuance. Then, as his inner monologue focused on the relevant thread, his fingers sprang into action.

When users, stakeholders, and even executives express a desire for “one place for all our data”, it is often more than a simple regurgitation of the misinterpretation of techno-wisdom. Many times they are genuinely expressing a need but, like a child, they lack the means to articulate their real problem. This presents the organization with a challenge. Because both purveyors and end users of technologies are preaching and praying, respectively, for the same thing it creates a false sense of alignment that leads to overconfident, premature action.

Suggesting to have “one place for all your data” is an expression of a solution without an originating problem. A true solution can only be a function of the actual problem. Instead of merely pacifying the customer by providing explicitly what they ask for, a wise solution provider should seek to understand the true needs of the customer, irrespective of their ability to succinctly express them at the beginning of a discovery process. Unfortunately, many technologists and consultants immediately jump into solutioning because the customer has uttered the magic words required for an infinite supply of meaningless billable hours, creating the appearance of value creation while avoiding the genuine challenge of defining a meaningful problem and objective. So, the operative question then emerges: what do they really mean when they ask for “one place for all of their data”?

  • They are asking for better knowledge of what data exists and where it resides.
  • They are asking for the means to use it performantly and cost-effectively.
  • They are asking for trustworthiness in their findings.
  • They are asking for a sense of confidence, certainty, and control.

Each of these underlying issues can be addressed without enterprise-wide centralization of data, or even much centralization at all. Furthermore, centralization of data can occur without actually addressing any of the above needs. The question then shifts again: “Is the centralization of data and data systems really the only weapon in our arsenal? Or are there other methods, perhaps better methods, that can help us achieve the actual desired outcomes better and faster?”

The Cost-Basis for Considering Data Centralization

Carl patiently awaited the codec agent’s acceptance of his latest contribution. He engrossed himself in an entertaining daydream (he preferred these to the doom-scrolling he had adopted with his former employer) until he was interrupted by the new text that appeared on his terminal.

>> Codec decompression algorithm accepted the latest submission. Simulator generated a 78% likelihood that a heuristic is still necessary for communicating this idea. 34% of divergent behaviors result in negative outcomes while the remaining 66% result in inaction. Is there a better way to express the techno-wisdom in question that doesn’t compromise the listener/observer?

Carl leaned back and exhaled. While some people would learn to ignore the falseness of the techno-wisdom, they would not take action when it was genuinely needed. This meant that the link between the input knowledge and the decompression assistance he had been coding wasn’t sufficiently self-evident yet. He needed to inject a patch that empowered as much as it protected. His hands floated millimeters above his keyboard as he considered his position and then began typing.

If we revisit the idea that when customers or stakeholders express a desire for “one place for all [their] data,” they are often genuinely signaling that there is a problem. But what precisely is that problem? The underlying ailments can be partially expressed as the four issues listed in the previous contribution. Each of these four issues has a common base class interpretation: they each represent a burden or a cost to using the data when left unaddressed.

Building on this notion, we can reframe the stakeholder request (and thereby the problem statement) from defining success based on the position or location of data to defining success based on the per-unit cost of using the data for an explicit, or even general, purpose. Therefore, a customer’s expressed desire for a centralized position or location of data is a proxy request for the true need for uniformly low-cost access to data across time.

What does this mean? For data access to be low-cost, every facet of its use should be as seamless and user-friendly as possible. It is a statement that is inclusive of all the underlying needs unexpressed by the customer.

Carl instinctively realized an example would be necessary here.

We’ll need to turn to an example: let’s consider a scenario with a group of users wishing to work with data. We can imagine many different types of data they want to interact with, all created by different business processes, managed by various owners, and stored in different locations, perhaps even with different database technologies—each with its unique characteristics. These characteristics then dictate the unique access-cost curve for each data modality. We can also imagine that all these users will interact with the data through some type of user interface.This interface could be purely point-and-click, an API, an export, a full-stack application, or a dashboard. It could be as a user or a developer. For the time being, we can keep the idea of the user interface in the abstract.

Now, conventional wisdom dictates that we need to put all of our data in “one place,” right? We’re going to presume that we’ve taken a literal interpretation of that and have physically located all the data in the same place, even forcing different modalities of data into the same database/same data system architecture.

The net effect of this is that while we have achieved the desired co-location of all the data, the ingest, and, in particular, the access-cost will vary greatly because we’ve likely implemented a least common denominator solution for storing and accessing the data. This means that while one type of data can be accessed performantly and cost-effectively, another type of data may suffer prohibitive difficulty to access, rendering it virtually unused—an infinite access-cost. Furthermore, the transplantation of data to a new system calls into question issues of ownership and data management. Separating the data from its source process or source system can introduce greater uncertainty than existed before, thereby increasing the access-cost. In summary, a literal centralization imposes wasteful limitations on diverse data modalities and introduces additional per-unit storage/access-costs!

Carl inserted a user story into the contribution to help drive home the example with something specific from the “real world”.

<USER_STORY_REF> <SOURCE: Large Electric Transmission Utility> I had been working with synchrophasor data since 2009, and up until 2017, the entirety of my work had been focused on real-time streaming applications and data systems, which, while challenging in its own right, doesn’t truly become a “big data” problem. The velocity of the data is quite fast, but the total volume at any point in time is not that large. The rest of the industry was no different; they too focused almost exclusively on the use of data in a real-time environment.Storage of historical data, while discussed, was an afterthought.

However, due to cultural challenges in integrating synchrophasor technology into the electric transmission control room, the prior eight years had failed to “close the loop” on the institutionalization of synchrophasor data. Evidently, the control room could not serve as the entry point for this new technology and modality of data into the organization. We required an entry point that could incubate, establish, and maintain a virtuous cycle of value creation and innovation at a faster pace. Therefore, we pivoted to using synchrophasor data as an engineering analysis tool rather than an operational tool.

Because this new mission mandates the ingest, storage, access, visualization, and analysis of large volumes of historical data, a new challenge emerged. Naturally, we considered our incumbent time-series “historian,” which held data from our SCADA systems, measured at 1Hz or less resolution. We mistakenly assumed it could also be used for time-series data at resolutions of 30Hz and above (e.g. synchrophasors). Vendors and consultants all said the same thing—the legacy historian would suffice and that many other people use it for synchrophasor data. Part of our willingness to believe stemmed from its successful use in the real-time environment (i.e. a “small” data problem). Of course, when this assumption met reality, it didn’t live up to its promise. It turns out that above the 1-5Hz sample rate, we end up with a very different computer science problem that requires a redesign of the entire ecosystem from the ground up. The solution became the adoption of a second data platform optimized for synchrophasors and engineering analysis of historical data. And that is the beginning of an entire story of its own.

On a related note, the common and frequent sales pitches for “network model management” are another instance where a utopic vision for the centralization of data ignores practical needs of consumers and undermines business objectives with distractions and false promises.</USER_STORY_REF> 

The correct interpretation of the referenced user story is that due to the substantial difference in the modality of sensor data, the “obvious” solution of co-locating the new data in the same system was invalid. In order to ensure proper performance and cost-effectiveness, a second platform optimized for synchrophasor data was required.Furthermore, it did not make sense to disrupt the existing platform and integrate into the new platform for the sake of unification. Such a decision could be delayed as technologies and use cases became better understood and developed. Therefore, a multi-platform solution was the right approach to maintaining a uniform access-cost for each modality of data.

So, what is a better way to think and talk about the uniformity that drives down access-cost? Let’s start again with our group of enthusiastic users and the same abstract user interface. Again, we have a diverse modality of data sets potentially disparately located. But rather than constraining them to the same physical location, we contain them in structured environments and expose them with technologies uniquely optimized for their cost-performance characteristics relative to the demand/utilization of that data. Then, we can create an abstract boundary around each of those data sources, which is not necessarily physical in nature. If we have done everything right, we end up with a uniform cost-performance curve for each data type relative to their demand/utilization, ensuring that the read/write of no single data type becomes the bottleneck for the user workflow. So, when we say to “put all of your data into one place,” we are setting an expectation for cost-performance of the utilization of that data, which, of course, also includes soft benefits such as knowledge of what data exists and other data governance benefits of co-locating data.

Carl’s terminal hummed while processing the latest entry, and he allowed himself to imagine that it was the sound of satisfaction of the algorithm, savoring each morsel as it consumed and metabolized his contribution. His fantasy was disrupted by the next message in his terminal. It appears that the algorithm still felt that there was a gap in its understanding.

>> Codec decompression updated successfully; Analysis shows that within session A274R08 the developer has only provided a critique of the existing wisdom and an alternative ideal to work towards surfaced in the form of a superior heuristic. However, this still neglects how one can move closer towards the ideal state incrementally therefore avoiding the perils of a single, large effort.

Carl felt exposed and embarrassed that his junior status as a developer was evident in his work. He tried to remind himself that this is why the agents use simulators and that development is iterative. Nobody sits down to a terminal and just starts coding and gets it right the first time through. Each submission, each contribution to the codec is, in essence, a hypothesis that must be tested by confronting it with reality (or at least the reality that we could simulate). Perhaps therein lies his solution to the issue that the codec agent surfaced. Embracing an ideal, even with data and data systems, is not a single hop from the status quo to the ideal state. It, too, is an iterative, evolutionary process. And while he had established the need for and philosophy of a better ideal, he hadn’t proposed a philosophy for guiding people and organizations towards that ideal.

A quick glance at the clock showed it was past time for him to head out for the day. He opened his notebook and quickly took down some notes so that he could pickup where he left off tomorrow.

He scribbled: 

Don’t allow yourself to conflate intentions with methodology...

He closed his notebook, removed the ticket card from the surface of the workstation and the screen slowly dimmed in response. He knew if he wanted to tackle this next challenge he would have to rest and recover properly. It was time for him to head home. He stored his ticket safely in a personal box, put on his jacket, grabbed his bag, and gave his mind permission to forget his latest challenge before he reached the sanctuary of his home.


Sign Up

Subscribe to receive email notifications whenever a new blog post is published.
You can unsubscribe at any time.


Kevin

This is a test

1 thought on “Codec Debug Session 2: Tell Us What You Really Mean”

Comments are closed.