Show HN:基于 ClickHouse 的 500 TB 互联网索引,采用拥堵定价
事件
Meet Scry, where I have a 500 TB NVMe internet index (I'm doing my best indexing and normalizing all the intelligence explosion alpha) that you can run ~arbitrary readonly SQL and some of Datalog over, and I handle the problem of resource-contention with congestion-based micro-auction pricing. When there's capacity, the service is free for non-commercial use. --- Hello. It's 2026, we're training simulated fruit fly brains to play Beat Saber, do we still have to be stuck with internet (re)search as fn: natural language -> black box we can't do anything about -> ranked_list/summary? There is a long history of people trying to do very fancy things that end up being done in relational databases and a little SQL. There is a gravity to them, a bitter lesson, just like scaling of generalized ml training methods. I mean many, many information products can be built off essentially giant real-time OLAP databases and frontier LLMs writing brilliant SQL+Datalog+vector+Jev etc. queries. Google Search, Tavily, Exa essentially have the problem of mapping your agents' context you are willing to provide, to a tiny subset of their index. You pay a fixed cost to an extremely hard problem that has a distribution of hardness, which means YOU eat the downsides when they are running out of budgeted compute to help you out. Their algorithms are opaque to the caller, there's really not much user control, and there's not a serious opportunity to communally
来源
本条目由采集管线自动抓取并发布,完整内容见下方来源链接。
*采集源:Hacker News*