5 comments

  • yashdotrv 23 hours ago
    Hi HN,

    This is Yash, founding team at Parseable (https://github.com/parseablehq).

    We've built an open source observability data lake using Rust, that handles high-cardinality data at around 100M time series in production (https://www.parseable.com/blog/how-parseable-handles-100-mil...)

    Our architecture is built around columnar design, and we use Apache Arrow for in-memory columnar processing and Apache Parquet for durable columnar storage on S3-compatible object storage. In Parseable, every labels stay as columns in the data instead of becoming a large long-lived per-series index like many TSDBs.

    Also, one thing we’ve been thinking about a lot is how observability changes as agents become part of day-to-day engineering workflows. They're not just another service, they produce traces, tool calls, prompts, intermediate decisions, errors, costs, and sometimes sensitive business context.

    Observing them matters just as much as observing any other system. But it is equally important to decide where that telemetry data should reside. Our view is that teams should be able to keep these observability data close to them: in their own object storage, under their own retention, access, and compliance controls.

    • goldeneye13_ 21 hours ago
      This looks super interesting. Question about the scale, I thought Thanos and some other Prometheus variants can handle about 100 million active time series. I would have expected your solution to scale to billions. Have you not pushed it past 100 million or am I maybe missing something.
      • kyxsc 15 hours ago
        I currently manage 50m active series and it’s brain dead easy using Mimir (which can also easily do 100-200m).

        That being said, perhaps this Parseable project is much more efficient (RAM) in storing that 1b samples compared to Mimir (RAM). Once we add in object stores which cheapen the storage cost by orders of magnitude (by going from memory to disk), the comparison is even weaker. Mimir and Cortex/Thanos are quite happy to pull cold data from S3.

        So I too expected to see “billions” as well.. hmm.

      • parmesant 21 hours ago
        We haven't yet tried pushing it to the scale of billions yet. The max that we've gone to is 150-180 million.
      • nikhil4usinha 21 hours ago
        Fair question. 100M isn't a ceiling, it's what we have seen in that deployment. We have not run a billion series test yet.

        The reason we think it scales differently - labels are just columns in Parquet, so there is no per series index that grows with cardinality. In that deployment one label alone has ~2.5M distinct values among 500+ labels, which would be painful for an index based TSDB but here is just a high cardinality column. What drives cost for us is ingestion rate (data points/s) and how much data a query has to scan for a particular time range not series count. Ingest scales horizontally by adding ingestors, and queries prune by time partition and column stats.

        A billion series benchmark is on our list, and we'll publish the numbers when we run it.

        • kyxsc 15 hours ago
          Definitely publish it! I’ll be following closely! 100M seems a bit too low to turn heads, but cool project nevertheless. Always exciting to see open source observability tools pop up!
    • msandford 22 hours ago
      100M active time series is good information, but what's the data rate for each time series it can handle? One update per minute or 10 per second? There's a factor of 600 difference there. Neither is obviously insanely the wrong update rate.
      • nikhil4usinha 21 hours ago
        scrape interval is 15s and sustained ingestion we have seen is ~3M samples/sec that is ~300 TB/day of raw ingest payload, when stored on object store as parquet, the data gets compressed to 99% which makes it 3 TB/day. The 100M figure is total unique series seen over time. For a sense of per metric cardinality, one metric that has the highest cardinality label (2.5 M distinct values) shows ~6M active series per hour.
    • codegeek 15 hours ago
      Your pricing page calculator is a bit strange. The minimum daily is set to 1 TB which is too high. Are you not interested in working with companies that have lesser ingestion ?
      • yashdotrv 7 hours ago
        Oh I see, what you say! ...but that’s intentional... we start the calculator at 1 TB/day because beyond this scale it's usually where observability pricing starts to hurt and ROI discussions kicks in. So, the point we’re trying to show is that, even at 1 TB/day, Parseable will still cost you lesser than others, and easier to operate too. Also the best part is that you will have the complete ownership of your data.
    • gustavohoa 20 hours ago
      How does the 100M active series deployment looks like? How many ingestors are there? What's each instance size? How big is the querier so that it can query across a metric with millions of active series?
      • parmesant 19 hours ago
        Sizing for this kind of deployment was a lot of fun! We went ahead with- 5x Ingestors, each with 64 vcpu 128 GB

        5x Queriers, each with 64 vcpu 192 GB

        The current utilization sits comfortably at 10-15 vcpu and 20-30 GB memory for the ingestors 20-40 vcpu and 40-60 GB memory for the queriers

        Ample of headroom for transient spikes and planned near-future growth

    • jiggawatts 15 hours ago
      > S3-compatible object storage

      Azure doesn't exist.

  • nwmcsween 20 hours ago
    Low level how does this compare to Victoriametrics, you can get the high cardinality with clickhouse and other columnar dbs but the tradeoff is more ram usage, slower queries, io, etc
    • parmesant 18 hours ago
      the tradeoff with regards to slower queries and io only comes into play if the TSDB is performing a narrow lookup on a handful of series. In that case it’ll be faster. But when you scale up to 100s of millions, columnar dbs like Parseable win because-

      a) there's no per-series inverted index and labels are parquet columns so memory is not bounded by cardinality

      b) data lives on much cheaper object storage (parseable gives an option to cache data locally to remove io bound latency)

      c) columnar store helps with faster data scanning by aggressively pruning and filtering data out

  • simonw 18 hours ago
    The open source version doesn't accept protobufs, but does accept JSON.

    I found this out because I set Codex the task of running this locally agains another of my apps and it worked around the limitation by running this proxy: https://gist.github.com/simonw/b0e61a0aa8e3f7d30f27ce2f747c9...

    ... but it turns out my stack can emit JSON just fine, so I switched to that instead. Here's me TIL write-up of getting Parseable running locally https://til.simonwillison.net/datasette/datasette-parseable-...

  • usernametaken29 19 hours ago
    How’s this any different then Iceberg?
    • pkb_98 7 hours ago
      Iceberg is a table format. It adds snapshots and a catalog layer on top of Parquet files. Parseable is a full-stack observability tool. Its storage layer could theoretically use Iceberg as the table format on top of Parquet file format. Iceberg is not a direct alternative to Parseable.
    • nylonstrung 17 hours ago
      I'm also struggling to understand what the novel tech here actually is

      What could I do with this that I couldn't achieve with iceberg-rs + DataFusion + parquet/vortex

      They repeatedly talk about "80% size reduction with compression. Isn't that essentially just the default parquet compression ratio?

      What exactly is unique here

      • kyxsc 15 hours ago
        I’m curious too. We need to see 1b+ numbers or else it just seems like we’re leveraging compression here, which is great but it’s hard to understand why it’s novel and why it’s the best solution no contest.
        • nylonstrung 1 hour ago
          Parquet is also 13 years old and not the GOAT at compression either, there's been a lot of developments in research since that like Fastlanes and Btrblocks which Vortex implements to put it at the Pareto frontier
        • usernametaken29 8 hours ago
          I can archive my data just fine by using an access based storage tiering like S3 - have your cake and eat it too. Gets around the headache of actually querying all of the 1B+ records because realistically there’s very few use cases for that. If you ever really need it, spin up Kafka and run through the entire archive.
    • parmesant 18 hours ago
      [flagged]
  • vivzkestrel 7 hours ago
    - comparison with timescaledb please?