14 Comments
User's avatar
Cyndy Bunn's avatar

I follow you on X as well. Thank you for explaining your work! I wish I had your talent.

Greg Tibbitts's avatar

DR - that is fantastic, thank you.

I love the Plain English about Google and web sites - that resonates. We need to get that to X and broadcast it.

Well done!

Rita's avatar

What I find so incomprehensible is that you had to do this in the first place. Making large datasets searchable is the first task, not the last, when creating a functional database. At least, that's how it was when I learned programming back when dinosaurs roamed the earth.

Aarati Martino's avatar

Thank you for writing this! It's crazy to think that a device we've had since forever -- the index at the back of a book -- is the key to harnessing all this information. That and the fact that you can store so much data in RAM on modern machines and just zip through it at will!

These (amongst others) are things we figured out at Google decades ago, it is great to see them applied here to government transparency. I hope you have plans to join with other public dbs like opensecrets or those at https://console.cloud.google.com/marketplace/browse?filter=solution-type:dataset&inv=1&invt=Abp9SA to discover more insights.

Dan Star's avatar

This is the best explanation of Search in databases that I have ever read. DOGE needs to modernize the Patent Office. My friend is an IP Paralegal and the Patent Office processes have room for big improvement.

Phil Turmel's avatar

Did you build your reverse index with PostgreSQL's native full-text indexing tools, or is it entirely custom? Or a mix?

DataRepublican's avatar

Entirely custom, and it runs entirely on the front end. There is no actual database engine or backend involved.

David Silverberg's avatar

DataR, may I ask you (1) the order of magnitude of the total data you need to index (2) whether the data is located on a centralized local device after initial data collection (prior to indexing), and last (3) can you give users database procedures (perhaps with a point-and-click UI) in addition to queries -- so users can do analysis. And if not, is that a potential desired upgrade capability?

I apologize if my question is much too presumptuous. If you would rather not say here, you can contact my office at Aestiva.com to get a hold of me.

Mark Windsor's avatar

This was helpful in broadly understanding what the spendthrift elites have been doing to cloak their crimes! Hiding their needles in thousands and thousands of haystacks. This resonates with what Nancy Pelosi, (as Speaker) said, "You will have to pass it to find out what is in it!" Not anymore with @DataRepublican indexing! LOL

Steven Westerberg's avatar

THANK YOU !!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!!

Penelope Zela & David Moksha's avatar

What is the role of DOGE in all this? What is the relationship between UsaSpending and DOGE and DataRepublican?

David Silverberg's avatar

As a tech company owner with his own DB architecture that includes standard record-based and word-based indexing I appreciate your thoughtful explanation. Of course, issues related to overcoming 32 and 64-bit issues, memory conservation during indexing, and the ability to perform joins without pre-indexing, are issues we had to overcome that you may have encountered.

More than your technological prowess however, I thank you for being a technologist with a strong moral compass. It would seem that is not asking for much, but apparently, it is.

Peter Starlite's avatar

We miss you on SubStack Jennica🫶