<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
  <channel>
    <title>Observability on App Coding</title>
    <link>https://appcoding.com/tags/observability/</link>
    <description>Recent content in Observability on App Coding</description>
    <generator>Hugo</generator>
    <language>en-us</language>
    <lastBuildDate>Mon, 05 Oct 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://appcoding.com/tags/observability/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>A Single-File Log Database: Pipe Logs In, Query Them With SQL, Hand the File to Anyone</title>
      <link>https://appcoding.com/a-single-file-log-database-pipe-logs-in-query-them-with-sql-hand-the-file-to-anyone/</link>
      <pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://appcoding.com/a-single-file-log-database-pipe-logs-in-query-them-with-sql-hand-the-file-to-anyone/</guid>
      <description>&lt;p&gt;The incident is over, the postmortem is Thursday, and someone wants the logs from the bad three hours. You can attach a gzip nobody will open, paste a screenshot of a grep, or stand up a search cluster to answer one question. The question is usually &amp;ldquo;which paths returned 500, and when did it start&amp;rdquo;, which has a &lt;code&gt;GROUP BY&lt;/code&gt; hiding in it, and a &lt;code&gt;GROUP BY&lt;/code&gt; is where grep runs out. A full &lt;a href=&#34;https://apicourse.com/api-observability-logging-metrics-and-distributed-tracing/&#34;&gt;observability stack&lt;/a&gt; answers it, but that&amp;rsquo;s a heavy answer for a team that doesn&amp;rsquo;t already run one.&lt;/p&gt;</description>
    </item>
    <item>
      <title>A Telemetry Database That Forgets on Purpose Keeps Rollups and Anomalies Instead of Raw Rows</title>
      <link>https://appcoding.com/a-telemetry-database-that-forgets-on-purpose-keeps-rollups-and-anomalies-instead-of-raw-rows/</link>
      <pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://appcoding.com/a-telemetry-database-that-forgets-on-purpose-keeps-rollups-and-anomalies-instead-of-raw-rows/</guid>
      <description>&lt;p&gt;At 300 requests a second, one latency metric adds about 26 million rows a day. Nobody will read most of them again. What people ask for a month later is the p99 per route for last Tuesday, and the one request that took nine seconds. A rule that deletes rows after 30 days throws away both answers. Keeping every row forever pays storage bills for data nobody reads.&lt;/p&gt;&#xA;&lt;p&gt;The telemetry store worth building decides at write time what has to survive. It folds everything else into coarser buckets on a schedule and keeps the few raw events that carry information, whole. Forgetting becomes part of the schema, declared in a policy, instead of a cron job that deletes by age. This is the argument for &lt;a href=&#34;https://apicoding.com/the-best-small-infrastructure-tools-reduce-data-near-the-source-instead-of-storing-more/&#34;&gt;reducing data near the source&lt;/a&gt;, applied to time series.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Debug From a Timeline of State Changes Instead of Ten Thousand Log Lines</title>
      <link>https://appcoding.com/debug-from-a-timeline-of-state-changes-instead-of-ten-thousand-log-lines/</link>
      <pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://appcoding.com/debug-from-a-timeline-of-state-changes-instead-of-ten-thousand-log-lines/</guid>
      <description>&lt;p&gt;A ticket comes in: order 8812 shows as cancelled, but the customer&amp;rsquo;s card was charged. You open log search. The checkout service wrote a couple of hundred lines for that request; the payment webhook handler wrote a few dozen; the job that expires unpaid orders wrote a line for every order it looked at that minute. The answer is in there somewhere. Forty minutes later you find it. The expiry job cancelled the order thirty seconds after checkout began, and the payment provider&amp;rsquo;s webhook landed half a second after that.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Group Errors by Behavioral Signature Instead of Stack Trace Text</title>
      <link>https://appcoding.com/group-errors-by-behavioral-signature-instead-of-stack-trace-text/</link>
      <pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://appcoding.com/group-errors-by-behavioral-signature-instead-of-stack-trace-text/</guid>
      <description>&lt;p&gt;Add six lines to the top of a file and ship it. Every stack trace through that file now carries different line numbers, so a tracker that hashes the trace files a bug it has seen a hundred times as a brand-new issue. Meanwhile an older issue titled &lt;code&gt;TimeoutError&lt;/code&gt; has collected comments from people who each fixed a different timeout and wondered why it never closed.&lt;/p&gt;&#xA;&lt;p&gt;Grouping fails in both directions at once. One bug turns into hundreds of issues because the message embeds an order ID, the line numbers moved, or a framework upgrade changed the frames in the middle. Two bugs turn into one because every timeout in the codebase raises the same exception type from the same retry helper.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Most Log Volume Is a Few Templates Repeated, So Reduce Logs Locally Before Shipping Them</title>
      <link>https://appcoding.com/most-log-volume-is-a-few-templates-repeated-so-reduce-logs-locally-before-shipping-them/</link>
      <pubDate>Mon, 05 Oct 2026 00:00:00 +0000</pubDate>
      <guid>https://appcoding.com/most-log-volume-is-a-few-templates-repeated-so-reduce-logs-locally-before-shipping-them/</guid>
      <description>&lt;p&gt;Take a day of production logs from one service, mask the numbers, IDs and quoted strings in every line, and sort the resulting patterns by how often they occur. The top of that list is short and dull: health checks, cache hits, a retry notice from a client library nobody owns. Each of those lines was serialized, shipped, parsed and indexed millions of times to say the same thing, and you paid for every copy. Run it on your own logs. The claim is easy to check, so don&amp;rsquo;t take it on faith.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
