The unsustainable cost of observability
When using cloud services, nothing is magic; every second of processing is monitored and paid for so that service availability is at its maximum possible threshold. To fulfill this promise, cloud providers or log storage services need their own Internal Telemetry for the services they offer for auditing, billing, and even security purposes. Furthermore, they must offer this service information to their customers as a vital functionality that cannot fail. This volume of event logs is stored without the proper preparation, since processing them as they are demands a disproportionate use of CPU and RAM—even for the hardware of these large companies—just to remain operational.
Fulfilling the operational promise to their customers in certain scenarios is even more expensive than the actual service the customer requested. This hidden work must be paid for, so that fixed monitoring margin is always reflected in the customers' bills month after month, a silent tax that covers up the underlying problem: observability.
Today, web systems use technologies that send a larger amount of logs than 10 years ago. These logs now have structure, but they are highly dynamic internally (Schema Drift). This structural change is a headache for log ingestion, and the CPU is the one that pays the price for this new nesting. The creation of the compressed file is another pain point in terms of RAM, since the file compression is hosted in memory, consuming considerable RAM (Memory buffering).
The industry has attempted to mitigate this chaos by relying on traditional engines based on inverted indices. However, this strategy has become an engineering trap for themselves. By attempting to tokenize and index every field of a constantly mutating JSON, the engine ends up building an index that often consumes more memory, CPU, and disk space than the original logs. Instead of solving the volume issue, the architecture simply duplicates the problem, forcing companies to pay for the storage of the log and its massive index.
As an escape route, many architects try to migrate towards classic columnar formats designed for analytics. While brilliant for traditional relational databases, these formats collapse when faced with the operational reality of observability. Their rigidity makes them inefficient and slow when dealing with cardinality and schemas that change frequently. Attempting to force dynamic data into rigid containers ends up fragmenting storage and skyrocketing compute costs during queries.
Finally, we can see that the true frontier of the log storage industry is to stop thinking only about scaling hardware to sustain an inefficient or legacy architecture, but rather to rethink the architecture of ingestion and storage.
El insostenible costo de la observabilidad
Cuando se usan servicios de la nube, nada es magia, cada segundo de procesamiento es vigilado y pagado para que la disponibilidad de los servicios esté en su máximo umbral posible. Para realizar esta promesa los proveedores de nube o servicios de almacenamiento de logs necesitan su propia Telemetría Interna de los servicios que ofrecen por temas de auditoría, facturación y hasta seguridad. Además, deben ofrecer la información de sus servicios a sus clientes como una funcionalidad vital y que no puede fallar. Este volumen de los registros de eventos se almacena sin la preparación correcta ya que procesarlos como están exige un uso desmesurado de CPU y RAM incluso para el hardware de estas grandes empresas solo para mantenerse operativos.
Cumplir la promesa de operación a sus clientes en ciertos escenarios es hasta más caro que el propio servicio que el cliente pidió. Este trabajo oculto debe pagarse, entonces ese margen fijo de monitoreo siempre se ve reflejado en las facturas de los clientes mes a mes, un impuesto silencioso que tapa el problema de fondo, la observabilidad.
Hoy en día los sistemas web usan tecnologías que envían más cantidad de logs que hace 10 años. Estos logs ahora tienen estructura altamente dinámica internamente (Schema Drift). Este cambio estructural es un dolor de cabeza para la ingesta de logs, y es la CPU quien paga el precio de este nuevo anidamiento. La creación del archivo comprimido es otro dolor en términos de RAM ya que la compresión del archivo se aloja en memoria consumiendo RAM considerable (Buffering de memoria).
La industria ha intentado mitigar este caos recurriendo a motores tradicionales basados en índices invertidos. Sin embargo, esta estrategia se ha convertido en una trampa de ingeniería para ellos mismos. Al intentar tokenizar e indexar cada campo de un JSON en constante mutación, el motor termina construyendo un índice que a menudo consume más memoria, CPU y disco que los logs originales. En lugar de resolver el volumen, la arquitectura simplemente duplica el problema, obligando a las empresas a pagar por el almacenamiento del log y por el de su propio índice masivo.
Como vía de escape, muchos arquitectos intentan migrar hacia formatos columnares clásicos diseñados para analítica. Aunque brillantes para bases de datos relacionales tradicionales, estos formatos colapsan frente a la realidad operativa de la observabilidad. Su rigidez los vuelve ineficientes y lentos cuando se enfrentan a la cardinalidad y a los esquemas que cambian frecuentemente. Intentar forzar datos dinámicos dentro de contenedores rígidos termina fragmentando el almacenamiento y disparando los costos de computación al momento de consultar.
Finalmente podemos ver que la verdadera frontera de la industria de almacenamiento de logs es dejar de pensar en solo escalar el hardware para sostener una arquitectura ineficiente o legacy, sino replantear la arquitectura de la ingesta y almacenamiento.