9 comments

  • sgt 0 minutes ago
    Cool project but .. reality is that people will generally not choose pgrust over Postgres, even 5-10 years from now. The problem is not that it may be technically superior and faster by then, it's that it's not built by the trusted Postgres team. There's a lot more to trust than development velocity or performance. It's also about the longevity and continuity of a critical piece of technology.
  • malisper 1 hour ago
    Author here. Let me know if you have any questions about the post or about pgrust.

    Let me take a shot at answering what I think will be the most common question: how can I trust pgrust? Our #1 priority right now is correctness. Over the past two weeks, I've done a mix of formal verification and differential fuzz testing. We've been able to prove over 1000 user facing functions have the exact same logic in both pgrust and postgres (see the proofs directory if you're curious). For cases where formal verification is not easy, we've taken the c implementation of a function and the rust implementation of a function and ran millions of inputs through each of them and confirmed they gave the same results every time.

    We've only covered about 15% of the surface area so far, but in the process, we've discovered ~100 bugs in pgrust and ~20 bugs in Postgres itself. My favorite postgres bug we found is this one[0]. Postgres has a quadtree implementation. Due to floating point rounding, it was possible for a point to be neither above, nor below, nor even with the center point of the quadtree.

    We've also entered engagements with Antithesis[1] to do Jepsen style fault testing and Aretta[2] to do more serious formal verification.

    If you want to support the project, the easiest way is to give us a star on GitHub[3]

    [0] https://www.postgresql.org/message-id/19597-39c532e61d78dff6...

    [1] https://antithesis.com/

    [2] https://aretta.ai/

    [3] https://github.com/malisper/pgrust

    • throwaway7783 35 minutes ago
      This is a great project. Thank you!

      A question on 20s postgresql time - It does not look like you are accounting for reading data from disk? Wouldn't the aggregation query have to load data from disk first? Or is it somewhat guaranteed that the table is already in memory? The Rust version is clearly in memory (I am no rust expert, so that may not even be actually in memory, if its a generator).

      • malisper 33 minutes ago
        > A question on 20s postgresql time - It does not look like you are accounting for reading data from disk

        I choose the data size so that it would fit in memory on the machine I was testing on. fwiw, there's still a ton of overhead Postgres has that the toy example does not. For example Postgres will serialize the numbers into tuples and need to deserialize them to execute the query. That's why it's not an apples-to-apples comparison

    • jnwatson 1 hour ago
      The floating point comparison bug is nightmare fuel. I could look at that for years and never spot the mistake.
      • bee_rider 1 hour ago
        On the bright side it could probably run for years without hitting the mistake as well. But it is nice to get it out of there.
    • btown 1 hour ago
      If someone wanted to use this as a real-time WAL-tracking read-only mirror of a live production database, for analytics work, is it ready for that use case yet?
      • malisper 1 hour ago
        You can try it. We're happy to help you with it, but expect there to be issues to work through. You would want to do it for something non-critical
    • lizimo 1 hour ago
      Is `pgrcolumnar` the default storage layout for tables? It would be cool if the same storage engine outperforms vanilla Postgres under both OLTP and OLAP workloads.

      AlloyDB from Google Cloud uses columnar storage like a secondary index, while the relations are still stored in TOAST.

      • malisper 1 hour ago
        pgrcolumnar is not the default storage method. Right now, it's exposed as a table access method. There's lots of design space for how to do this so I want to avoid pre-committing to anything
    • andriy_koval 52 minutes ago
      what is your vision of this project? Do you think pgrust will eventually be prod ready?
      • malisper 47 minutes ago
        I want to build the best database possible. While Postgres is great, there are a lot of core issues that have been around for over a decade. We're working hard to get pgrust production-ready, and it will definitely be production-ready in the near future. I wouldn't be putting hundreds of thousands of dollars into this project if I didn't think we could build a production-ready database.
    • doctorpangloss 1 hour ago
      “Show me the prompt.”
  • AsyncBanana 24 minutes ago
    You have no idea how long I have been waiting for adaptive planning. One of my biggest annoyances with the Postgres core team has been their reluctance to implement any sort of adaptive planning despite it, at this point, being a well-established technique that has been implemented in multiple production databases. I hope this, at the very least, proves the viability of this model outside of academic/niche contexts.
  • rastignack 1 hour ago
    I would be interested about a more detailed architecture overview of the io scheduler (like this: https://www.scylladb.com/2021/04/06/scyllas-new-io-scheduler...) and the thread scheduler.

    PostgreSQL has historically been bad at managing the noisy neighbor problem, but with thread pools, and io priorities, it can be solved.

    Has this been tackled here ?

  • KolmogorovComp 12 minutes ago
    Thanks to the authors for choosing a license that respect users freedom, on top of being an awesome technical project.
  • Lucasoato 1 hour ago
    I’m curious to see how this compares to pgColumnar or other OLAP extensions.
  • cognitiveinline 3 hours ago
    pgrust seems to have good momentum. AGPL is an odd license for a non web project. Postgres is MIT-like, and that drove it's adoption.

    Have pgrust folks reconsidered this? Else, IMO we can have an independant rust port of pgrust, which can be MIT, which will garner more attention.

    • malisper 2 hours ago
      Author here. At least for databases, AGPL (or stricter) has become standard. The issue is it's so easy for megacorps (Amazon, Google, etc) to take a permissively licensed product and monetize it at the expense of the original standard.

      For instance, Mongo, Cockroach, and Materialize have all gone source available. We picked AGPL because it's the best balance between open source and prevents Amazon from just repackaging it and selling it.

      If AGPL is an issue for anyone, we would be happy to dual-license under a commercial license.

      • mey 1 hour ago
        There is nothing wrong with wanting to be compensated for your work, but for people like myself which use a cloud managed DB solution (GCP CloudSQL PostgreSQL) it means something like this would never be available.

        I consider AGPL a poison pill in my work. That is not true with a suitable commercial license, although I expect a lot more commercial product (support/features/etc). As you note, your objective is to prevent commercialization of your software, but radically speeding up analytics is primarily a concern of large organizations so it seems like a mismatch in purpose.

      • cognitiveinline 1 hour ago
        Sure, that's your prerogative, and kudos for not talking up open source. I'm not amazon size so can't use it, and AGPL is a no go for DB, don't want to be forced to open source my app because I use this!

        Will await a MIT based fork myself.

        • jnwatson 1 hour ago
          Why would AGPL force you to open source your app? Unless you literally compile your app with pgrust by modifying the pgrust source code, you're safe. Clients aren't bound by the AGPL because they aren't derived works.
          • xyzzy_plugh 1 hour ago
            Here we go again.

            AGPL is untested in courts. There is no definitive definition of what could be considered within the blast radius such that it would require AGPL licensing.

            There's a reason AGPL is banned at Google and most sane companies. It's simply too dangerous.

            You can't simply say "clients aren't bound" because it depends.

            I'd rather see the BSL used here to be perfectly honest. At least it's simple.

    • appplication 2 hours ago
      Oh bummer. I was really excited about pgrust but AGPL is a dealbreaker. Not for me personally, but it will never see wide adoption because it’s a banned license in most corporate environments. Lack of path to wide adoption means it’s dead in the water.

      It’s weird because those who actually care about optimized pg gains are most likely large corporate customers. Why make a product targeting them and license it in such a way they’ll never use it?

      This also hard blocks upstreaming any beneficial features into core Postgres.

      • karlmush 32 minutes ago
        AGPL seems like the right choice to me. I’m tired of companies like PlanetScale taking PostgreSQL, building a business on top of it, and then acting like PostgreSQL is theirs to control.
        • samlambert 18 minutes ago
          we have not once claimed postgres is under our control. i don't think you understand how open source works but thats ok.
      • guenthert 1 hour ago
        If it is indeed 300* faster, I'm sure more rational corporations will rethink their license policy or be left in the dust.
        • jacquesm 48 minutes ago
          License policies are made by lawyers not by programmers. And their competition will be in the exact same boat (different lawyers though). AGPL is so toxic that it tends to be checked for during M&A processes so even if the current batch of lawyers is ok with it there is a chance that a later batch of lawyers is not. Given that the target audience for this project is the larger companies you are going to end up with a very nice project and zero actual users or you will end up with AWS et all stealing your work. Databases are very hard to do successfully commercially, at a minimum you should dual license them (AGPL for 'home' use and commercial licensing for parties that will want to buy the upside but they'll demand support and other stuff besides).
          • hunterpayne 34 minutes ago
            This is the case when you have your own datacenter. This isn't as big a problem in the cloud. There are ways to write licenses that prevent cloud providers from stealing while allowing customers to use the software and being required to pay for it. The problem with the AGPL has to with its viral nature, not its provisions to prevent cloud vendor theft.
        • hunterpayne 41 minutes ago
          Its 300x faster for certain tests. I could pretty easily craft tests that do this on two different systems. The author mentioned this when they talked about being able to fit an entire ResultSet into memory. That's the real trick with performance. Very few workloads are CPU bound anymore (linear algebra on the CPU for example). Almost all workloads are memory bound. So its all about moving data from memory to network, back to memory and back to network, over and over again through your microservices or DBs. If the entire working set can fit in memory, you get at least a 10x performance boost. If you have to keep even a part of the working set on disk, its a huge performance loss. And the larger fraction of the working set on disk, the worse the performance loss.

          PS Learn how DBs do joins for more information. Specifically the differences between hash joins, merge joins and nested loop joins. They are basically fancy ways to page part of your working set to disk at huge performance penalties.

          PPS As memory gets more expensive, these techniques get more valuable. When it gets cheap, they lose value.

        • xyzzy_plugh 1 hour ago
          They could simply spend a few months and a few million tokens and get their own port, no?

          I doubt even 30000x faster would prompt a policy change.

          • hunterpayne 46 minutes ago
            No. This is system code. You let an LLM loose on it, it probably fixed 20 bugs and introduces 200 more plus 5 different performance regressions. In all fairness, your average app programmer would have the same problems. That's why it takes so long to learn to be a system programmer and why it takes so long to do anything on a systems source base. For reference, systems are OSes, DBs and compilers (although compilers are very different in many ways).

            Also, you can successfully sell a systems project that is only 10% faster. 30000x faster and they are throwing illegal and debauched things through your window to get access to your improvements.

            • skinfaxi 5 minutes ago
              > No. This is system code. You let an LLM loose on it, it probably fixed 20 bugs and introduces 200 more plus 5 different performance regressions.

              What does it being system code have to do with anything?

          • superb_dev 53 minutes ago
            Are you suggesting the AI just rewrites the whole thing under a different license? There’s no way that’s not more dicey than the AGPL license.
            • jacquesm 48 minutes ago
              That's exactly what they did here, I don't see the difference.
              • hunterpayne 27 minutes ago
                This is the difference. This guy took an existing source base, had Claude find specific bugs, then had Claude fix a specific bug which was then reviewed by a person. We also don't know if these changes introduce new problems yet. You are suggesting letting Claude write an entirely new source base. That's light-years away from what happened here.
              • yifanl 33 minutes ago
                They did it on permissively licensed code would be the difference.
        • kornelijus 20 minutes ago
          So, keyword 'rational', I'm not sure any sufficiently large company is a rational actor.

          Yes, at [tech corp dayjob], any dependency is likely to be banned for arbitrary reasons if you bring it to the attention of the wrong people. It doesn't have to go against any of our policies e.g. don't mention anything with GPL in the name around the "risk" people. In fact, do not ever talk to the "risk" people and hope they don't talk to you.

          Latest news: Apparently, devtools are a legal risk. Basic reverse-engineering of client-side JS is now banned.

          The delusions really seem to scale with headcount.

      • ForHackernews 33 minutes ago
        > It’s weird because those who actually care about optimized pg gains are most likely large corporate customers. Why make a product targeting them and license it in such a way they’ll never use it?

        Wait, sorry, you're asking why make something enterprise customers might pay for, and then not give it away to them for free?

    • ognarb 3 hours ago
      2 commits in the repo both generated by claude. This is AI slop, I wonder where you see good momentum?
      • Whitespace 3 hours ago
        main indeed has two commits, but it clearly states the location of the rest of the commits, so I wouldn't be critical of main itself.

          hey claude, do a breakthrough
          You can find the actual git history at the v0.2 github tag.
          
          Co-Authored-By: Fable <noreply@anthropic.com>
        
        Now we see https://github.com/malisper/pgrust/tree/v0.2 has almost 6000 commits in it, with the very first one on 2026-07-02. That's a lot of token momentum!

        It's easy to claim AI slop nowadays, but you should still mistrust-but-verify.

        • f311a 1 hour ago
          What's the reason for it? Does not make a lot of sense to keep all the commits elsewhere
          • malisper 22 minutes ago
            It's a reference to the prompt that found a counterexample to the Dinitz-Garg-Goemans conjecture

            > "do a breakthrough and find a structured counterexample"

      • esafak 3 hours ago
        4K stargazers in a week.
    • busterarm 2 hours ago
      Everything around Rust is political, so the license choices are also about political statement.
      • johnsonjo 2 hours ago
        Most official Rust projects are dual MIT/Apache licensed by convention [1] (and most Rust libraries from third parties I've seen that are open source MIT follow the MIT/Apache dual license), so seems like this library shouldn't just be AGPL for a typical political choice of a Rustacean?

        [1]: https://rust-lang.org/policies/licenses/

        • busterarm 2 hours ago
          official projects are usually run by sensible people who want to do things and aren't leading with ideology.

          We're literally talking about an "X but in Rust" project already...

  • Natalia724 33 minutes ago
    [dead]
  • xyzzy_plugh 1 hour ago
    [dead]