Kenny Meneses

KENNY MENESES

← Back to all posts

The unsustainable cost of observability

When using cloud services, nothing is magic; every second of processing is monitored and paid for so that service availability is at its maximum possible threshold. To fulfill this promise, cloud providers or log storage services need their own Internal Telemetry for the services they offer for auditing, billing, and even security purposes. Furthermore, they must offer this service information to their customers as a vital functionality that cannot fail. This volume of event logs is stored without the proper preparation, since processing them as they are demands a disproportionate use of CPU and RAM—even for the hardware of these large companies—just to remain operational.

Fulfilling the operational promise to their customers in certain scenarios is even more expensive than the actual service the customer requested. This hidden work must be paid for, so that fixed monitoring margin is always reflected in the customers' bills month after month, a silent tax that covers up the underlying problem: observability.

Today, web systems use technologies that send a larger amount of logs than 10 years ago. These logs now have structure, but they are highly dynamic internally (Schema Drift). This structural change is a headache for log ingestion, and the CPU is the one that pays the price for this new nesting. The creation of the compressed file is another pain point in terms of RAM, since the file compression is hosted in memory, consuming considerable RAM (Memory buffering).

The industry has attempted to mitigate this chaos by relying on traditional engines based on inverted indices. However, this strategy has become an engineering trap for themselves. By attempting to tokenize and index every field of a constantly mutating JSON, the engine ends up building an index that often consumes more memory, CPU, and disk space than the original logs. Instead of solving the volume issue, the architecture simply duplicates the problem, forcing companies to pay for the storage of the log and its massive index.

As an escape route, many architects try to migrate towards classic columnar formats designed for analytics. While brilliant for traditional relational databases, these formats collapse when faced with the operational reality of observability. Their rigidity makes them inefficient and slow when dealing with cardinality and schemas that change frequently. Attempting to force dynamic data into rigid containers ends up fragmenting storage and skyrocketing compute costs during queries.

Finally, we can see that the true frontier of the log storage industry is to stop thinking only about scaling hardware to sustain an inefficient or legacy architecture, but rather to rethink the architecture of ingestion and storage.