✍️ AI 改写 🔎 Meme筛选 🎯 Meme 猎手 🏆 Meme大神 🖼️ 梗图
2816 展示 > 赞 >
排序:
Tom Dörr AI @tom_doerr · 2026年07月28日 06:39
PwnPad 通过性价比极高的 PCB 挑战带你玩转硬件黑客,涵盖了 UART 分析、固件提取和侧信道攻击等硬核内容。

https://github.com/twelvesec/PwnPad https://x.com/tom_doerr/status/2081992953237590488
PwnPad teaches hardware hacking through affordable PCB challenges covering UART analysis, firmware extraction, and side-channel attacks.

https://github.com/twelvesec/PwnPad https://x.com/tom_doerr/status/2081992953237590488
图片
1 25 10.3K
Tom Dörr AI @tom_doerr · 2026年07月28日 07:08
A faithful remake of Super Mario Bros. featuring a full custom level editor, resource pack support, and custom characters. It includes recreations of the original game, The Lost Levels, and Super Mario Bros. Special.

https://github.com/JHDev2006/Super-Mario-Bros.-Remastered-Public https://x.com/tom_doerr/status/2082000304501850607
图片
0 16 10.2K
Neo Kim AI @systemdesignone · 2026年07月28日 16:33
程序员专供

你每天必用的 AI 编程神器是哪一个?
SOFTWARE ENGINEERS ONLY

What is one AI coding tool you use daily?
105 1 19.3K
karcharodon 加密 @kar4arodon · 2026年07月27日 12:06
这家伙一次预测都没做,却在 @Polymarket 上赚了 40 多万美元。

他的主页上完全没有动态,但反转来了:

他通过 Polymarket perps 赚了 40 多万美元。

这位大佬是 Polymarket 的顶级用户,靠 Polymarket perps 盈利超 40 万美元。

> 无关政治
> 无关加密
> 无关体育
√ 只有 Polymarket perps
This guys has made 0 predictions but still made +$400,000 on @Polymarket

There is no activity on his profile but here's a twist: come

He had over +$400,000 on Polymarket through Polymarket perps

This guy is a top 1 user of polymarket who made +$400,000 profit from polymarket perps

> There is no politics
> There is no crypto
> There is no sports
√ Only polymarket perps
21 1 10.1K
Sac AI @Saccc_c · 2026年07月28日 08:51
这就是我为什么要再一次订阅 claude 的原因

同一个prompt,同一个skill, codex 就是搓不出来这么高级的 3d 模型

使用的skill仓库:
https://github.com/img2threejs/img2threejs https://x.com/Saccc_c/status/2082026141620346915
23 11 10.9K
Adam也叫吉米 AI @Adam38363368936 · 2026年07月27日 09:30
继续写实感系列-认真工作的人最美
真实手机抓拍+年轻成年女性+具体生活动作,加入更明确的小事件、环境互动和轻微反差,真实感会更强

提示词:
竖版3:4,单张真实手机夜间抓拍,不要拼贴,不要四宫格,不要文字排版。

仅参考此前系列中年轻成年亚洲女性的五官、脸型、自然肤质和黑色长直发,保持人物外貌一致,只参考外貌,完全重建新场景。

深夜,一位年轻成年亚洲女性站在打印店内,穿浅灰色无袖针织上衣、深灰色宽松长裙和平底鞋,浅色帆布包放在脚边。她身体微微前倾,一只手扶着打印机边缘,另一只手接住刚刚打印出来的纸张,低头认真检查内容。

打印机旁边散放着订书机、透明文件袋、几张废纸和半杯没喝完的饮料。店内没有工作人员,只有冷白色顶灯,玻璃门外是安静的深夜街道。人物神情专注、有一点疲惫,不看镜头,不微笑。

普通手机后置摄像头,从玻璃门外或店内斜后方拍摄,中景构图,玻璃有轻微反光,画面略微倾斜,保留手机夜拍噪点和生活杂乱感。

负面提示词:办公室广告,豪华打印店,回头笑,正面对镜头,商业摆拍,纸张漂浮,打印机结构错误,多余手指,文字内容清晰可辨,塑料皮肤,水印。
图片图片
47 5 15.9K
娜美知识库 其他 @fhwofjow51260 · 2026年07月28日 13:12
图片图片图片
8 13 10.2K
周览资源 其他 @grgerwcwetwet · 2026年07月27日 10:58
💰重磅!价值6998的李一舟全套课程合集(目前全网已下架)89.5GB

链接:https://pan.quark.cn/s/4d7650c1b2e5 https://x.com/grgerwcwetwet/status/2081695777420918833
图片
25 22 10.1K
周览资源 其他 @grgerwcwetwet · 2026年07月28日 02:33
搞量化交易,别再东拼西凑找资料了。

推荐一个开源项目 Awesome Systematic Trading,几乎把量化交易的学习资源一次性整理齐了。

里面收录了 97 个量化库和工具、40+ 套经典策略、55 本量化书籍、23 个视频课程,从数据获取、回测、指标、风控到实盘框架,基本都能找到。

无论你是刚入门,还是想开发自己的量化策略,这个仓库都值得收藏,能少踩不少坑。

🔗 https://github.com/paperswithbacktest/awesome-systematic-trading
图片
2 55 11.1K
Fang 其他 @FLMdongtianfudi · 2026年07月28日 05:54
“有哪本书,你恨不得把它全部内容都背诵下来?”

答案点赞人数最多的竟然是这本:《Science Research Writing for non-native speakers of English》

立即保存👇
https://pan.quark.cn/s/b4c6fcc0f42b https://x.com/FLMdongtianfudi/status/2081981574095175894
图片
28 49 10.4K
铁锤人 AI @lxfater · 2026年07月28日 09:17
直接上传PDF给AI会让你多花4倍的token!!

大家可能以为, 把PDF传上到AI,和自己复制粘贴文本到对话框效果一样。但事实上,AI会用贵4倍的视觉能力看一遍你上传的文件。

但为什么AI要浪费4倍TOKEN方式处理文档呢?

因为这么处理能保留排版,表格这些信息, 回答质量会比所有同行高那么一点点。即使是高一点点也意味着第一,而第一永远享受最大的资源。

我们可以用下面这个7.5w star的开源项目,解决这个问题。

它能把 PDF、图片、Office 文档整篇变成 Markdown格式,同时不失去,关键的图像,排版的信息。 这样处理后,直接省下80%的TOKEN!!

之前这种服务,要么按页收费,要么要自己捣鼓配置,很难上手。

有了这个项目后,直接打开网页就能无限使用!!

关注我,每天分享一点实用AI技巧!!

https://github.com/opendatalab/mineru
17 27 20.8K
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在Robinhood链上花$2.1K买入大金狗 $YOLO,做到了 46x 回报,狂赚 $95.0K(约68.4万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/robinhood/address/C7KoyPop_0xfefd188448947c97e52774cd30d1cece5f201f17
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 19:01
发现一个牛B的钱包地址,在Robinhood链上花$96买入大金狗 $ASTEROID,做到了 151x 回报,狂赚 $14.5K(约10.5万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/robinhood/address/C7KoyPop_0xc3c52095cb29c66b15fbfcb5872acd3b8146b200
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在Robinhood链上花$877买入大金狗 $AI,做到了 123x 回报,狂赚 $108.0K(约77.8万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/robinhood/address/C7KoyPop_0xf23817f6d701f514861fe44371756e3226de8fef
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在Robinhood链上花$240买入大金狗 $GME,做到了 170x 回报,狂赚 $40.7K(约29.3万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/robinhood/address/C7KoyPop_0x729f4648090d8c6cadf3867e5c11ae8ec6e83af9
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在Robinhood链上花$63买入大金狗 $PONS,做到了 3156x 回报,狂赚 $198.1K(约142.6万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/robinhood/address/C7KoyPop_0x6099ff02f63c69162175249c2700c84c0a94dafa
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在Base链上花$4.4K买入大金狗 $BRIAN,做到了 52x 回报,狂赚 $232.4K(约167.3万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/base/address/C7KoyPop_0x5a0d8fcbdba47c78772dcf27753e7614cfd35e52
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 16:01
发现一个牛B的钱包地址,在BSC链上花$304买入大金狗 $MarsCoin,做到了 441x 回报,狂赚 $134.1K(约96.6万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/bsc/address/C7KoyPop_0xd2abcef40a51c779c2a890dd40f041ab995ed3af
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在SOL链上花$149买入大金狗 $旺旺,做到了 184x 回报,狂赚 $27.5K(约19.8万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/sol/address/C7KoyPop_DEnRzXaUFcdJh2jPgkbB3imE9j3EZj4w7ehKyiHTYSSD
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 01:01
发现一个牛B的钱包地址,在SOL链上花$68买入大金狗 $EVE,做到了 1129x 回报,狂赚 $76.7K(约55.2万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/sol/address/C7KoyPop_3GW928gZR6XZqmtjaZxnYpddCpP6HWQM82QjjsU9wsbT
0 0 0
Meme大神 其他 @meme_god · 2026年07月28日 19:01
发现一个牛B的钱包地址,在SOL链上花$111买入大金狗 $FRANK,做到了 306x 回报,狂赚 $33.9K(约24.4万人民币)!

跟着大神找金狗 👀,他的钱包地址:
https://gmgn.ai/sol/address/C7KoyPop_3GkGNBfyEp3GHSyRue5zuQZqp2Q28GovRqvSzDX7abgB
0 0 0
Nav Toor AI @heynavtoor · 2026年07月28日 10:07
运营商送的路由器,其实并不是你的。

它记录了你的每一次 DNS 查询、连接的每一个设备、你什么时候上线、又是什么时候下线。

2017 年 4 月 3 日,特朗普签署了 S.J.Res 34 法案。从此,运营商在法律上被允许在未经你同意的情况下出售你的浏览历史。

仅 2024 年上半年,AT&T 就收到了 152,561 次针对客户数据的执法调取要求,Verizon 收到了 145,195 次。数据直接来自各家公司的透明度报告。

Comcast 把盒子租给你,Xfinity xFi 网关:每月 15 美元。一年就是 180 美元,而且这硬件你永远不真正拥有。

Netgear 把基础安全功能锁在 Armor 会员里:每年 99.99 美元。

Amazon 把家长控制锁在 eero Plus 里:每月 9.99 美元。

而且大多数消费级路由器厂商在 3 到 5 年内就会停止固件支持。

接着在 2024 年 12 月,美国商务部对美国出货量最大的路由器品牌 TP-Link 展开调查,理由是可能与中国存在国家安全关联。

2026 年了,你竟然还在为访问自己的互联网而支付租金。

现在,来认识一下 OpenWrt。

一个免费开源的 Linux 操作系统,可刷入超过 2500 种路由器型号,让你夺回对网络的完全控制权。

它诞生于 2004 年 1 月,起因是 Linksys 出货的 WRT54G 运行了 Linux 固件。在 GPL 协议下,Linksys 必须公布源码。一个小社区拿到了代码,重新定义了家庭路由器的上限。

22 年过去了,OpenWrt 依然运行在 TP-Link、Netgear、Linksys、D-Link、Asus、GL.iNet、Ubiquiti、Xiaomi 上。那些厂商早已停更的老路由器,今天依然能用上最新的固件版本。

27,681 stars. 12,693 forks. GPL-2.0.

- 完整的 Linux root 权限,支持 8000+ 软件包
- 内置 WireGuard 和 OpenVPN 服务器
- 全局广告拦截和 DNS 过滤
- 访客网络、VLAN、针对每台设备的防火墙规则
- 针对视频通话的 SQM 流量整形
- 支持 Home Assistant、Prometheus、Grafana 导出
- LuCI Web 界面,无需命令行操作

OpenWrt 的成本是:

永远 0 元。支持你刷入的每一台设备。

Comcast 租路由器收 15 刀/月,OpenWrt 不收。
Netgear 安全服务收 99.99 刀/年,OpenWrt 不收。
eero 家长控制收 9.99 刀/月,OpenWrt 不收。

对 Xfinity 用户来说,每年能省 180 刀。
对 Netgear Orbi 用户来说,每年能省 99.99 刀。
对 eero 家庭来说,每年能省 120 刀。

你的路由器。你的固件。你的网络。

100% 开源。(链接在评论区)
Your ISP gave you a router.

It is not your router.

It can log every DNS query. Every device that connects. What time you go online. What time you stop.

On April 3, 2017, Trump signed S.J.Res 34 into law. In one signature, ISPs became legally allowed to sell your browsing history without your consent.

In the first half of 2024 alone, AT&T received 152,561 law enforcement demands for customer data. Verizon received 145,195. Straight from each company's own transparency report.

Then Comcast rents you the box. Xfinity xFi Gateway: $15 a month. $180 a year, for hardware you never own.

Then Netgear locks basic security behind Armor: $99.99 a year.

Then Amazon locks parental controls behind eero Plus: $9.99 a month.

Then most consumer router vendors drop firmware support within 3 to 5 years.

Then in December 2024, the U.S. Commerce Department opened an investigation into TP-Link, the most-shipped router brand in America, for potential national security ties to China.

You are renting access to your own internet in 2026.

Now meet OpenWrt.

A free and open source Linux operating system that flashes onto over 2,500 router models and gives you back full control of your own network.

Born in January 2004 after Linksys shipped the WRT54G running Linux firmware. Under the GPL, Linksys was required to publish the source. A small community took that code and rebuilt what a home router could be.

Twenty-two years later, OpenWrt runs on TP-Link. Netgear. Linksys. D-Link. Asus. GL.iNet. Ubiquiti. Xiaomi. Old routers whose vendors stopped patching still get fresh builds today.

27,681 stars. 12,693 forks. GPL-2.0.

- Full Linux root with 8,000+ packages
- WireGuard and OpenVPN server built in
- Ad blocking and DNS filtering network-wide
- Guest networks, VLANs, per-device firewall rules
- SQM traffic shaping for video calls
- Home Assistant, Prometheus, Grafana exporters
- LuCI web interface, no command line

Here is what OpenWrt costs:

Zero. Forever. On every device you flash.

Comcast charges $15 a month to rent a router. OpenWrt doesn't.
Netgear charges $99.99 a year for security. OpenWrt doesn't.
eero locks parental controls behind $9.99 a month. OpenWrt doesn't.

For one Xfinity subscriber, you save $180 a year.
For one Netgear Orbi owner, you save $99.99 a year.
For one eero household, you save $120 a year.

Your router. Your firmware. Your network.

100% Open Source. (Link in the comments)
图片
18 94 19.7K
Tom Dörr AI @tom_doerr · 2026年07月28日 03:59
使用支持 Keras 3 和 ONNX 的快速轻量级 OCR 模型,从图片中提取车牌号及地区。

https://github.com/ankandrew/fast-plate-ocr https://x.com/tom_doerr/status/2081952625684054285
Extracts license plate text and region from images using fast, lightweight OCR models supporting Keras 3 and ONNX.

https://github.com/ankandrew/fast-plate-ocr https://x.com/tom_doerr/status/2081952625684054285
图片
2 14 11.7K
Nav Toor AI @heynavtoor · 2026年07月28日 13:30
你给新员工配了台 $2,000 的 MacBook Pro。

然后每年还要花 $180 租 Microsoft 365 席位,Adobe $60,Slack $144,Zoom $240。所有这些开支全都绑在这一台电脑、一个员工、一张办公桌上。

接着 CFO 问你,公司这 400 台笔记本里到底有多少在用?都在谁手里?

你打开 Google 表格,发现一半的行都是空的。离职半年的员工名下居然还挂着公司资产。

于是 IT 找了 ServiceNow。ITAM 标准版每月每个用户 $100,专业版 $150,企业版 $200+。在 Vendr 上,ServiceNow 合同的中位数是每年 $124,364。实施费用更是授权费的 3 到 5 倍。

接着 IT 找了 Lansweeper。2,000 个资产每月要 $239。

再找 Freshservice。每个坐席 $19,超过 500 个资产后按量计费。

都 2026 年了,你居然要每年付六位数的美元,就为了搞清楚“DELL-4471 笔记本在谁手里”。

现在来看看 Snipe-IT。

一个免费开源的 IT 资产管理平台,部署在自己的服务器上,能追踪公司所有的笔记本、手机、软件授权和线缆。无云端限制,无代理,无席位费。

2013 年由开发者 Alison Gianotto 编写,当时她是纽约一家广告公司的 CTO,正好赶上办公室搬迁。她的团队当时用 Google 文档管着几百台笔记本。她试遍了所有的商业和开源工具,没一个好用的,于是干脆自己写了一个。

等她做完第一次完整的资产盘点,发现公司竟然丢了价值 $25,000 的资产。

她丈夫念叨了好几个月让她加个 PayPal 按钮。她最后照做了。结果两天后就有了付费客户。但 Snipe-IT 本身依然免费。

14,113 颗星,3,886 次 Fork。AGPL-3.0 协议。今天刚更新代码。

Snipe-IT 能给你提供:

- 一站式追踪所有笔记本、手机、显示器、线缆和配件
- 将资产分配给用户、部门或地点,支持完整历史记录
- 软件授权管理,支持席位追踪和过期提醒
- 自动生成条形码和二维码物理标签
- 自定义字段、状态和分类
- 支持 LDAP, SAML 和 Google Workspace SSO
- 所有接口都支持 REST API
- 能运行在任何 Linux 机器、Docker 或 $5 的 VPS 上

Snipe-IT 的价格:

零。永久免费。不收用户费,资产数量没上限。

ServiceNow 每个用户每月收 $100,Snipe-IT 不收。
Lansweeper 2,000 个资产每月收 $239,Snipe-IT 不收。
Freshservice 超过 500 个资产就计费,Snipe-IT 不收。
Ivanti 和 Flexera 把价格藏在销售电话后面,Snipe-IT 不搞这一套。

最离谱的是:

索尼、耶鲁大学、富士通,还有成千上万个付不起 $124,000 ServiceNow 账单的学区、医院和政府机构,都在用 Snipe-IT。Alison 在 Grokability 的官方头衔是“莫霍克酋长”(Chief Mohawk Officer)。

对于一个 50 人的 IT 团队,每年能省下 $60,000。
对于一家 500 人的公司,每年能省下 $600,000。

你的资产,你的数据库,你的服务器。

100% 开源。(链接在评论区)
You spent $2,000 on a MacBook Pro for a new hire.

Then $180 a year on the Microsoft 365 seat, $60 on Adobe, $144 on Slack, $240 on Zoom. All assigned to one laptop, one employee, one desk.

Then your CFO asked how many of the 400 laptops in this company are actually being used, and by whom.

You opened the Google Sheet. Half the rows were blank. Employees who left six months ago still had assets assigned.

So IT called ServiceNow. ITAM Standard is $100 per user per month. Professional $150. Enterprise $200+. The median ServiceNow contract on Vendr is $124,364 a year. Implementation runs three to five times the license fee.

Then IT called Lansweeper. $239 a month for 2,000 assets.

Then IT called Freshservice. $19 per agent, then metered on every asset over 500.

You are paying six figures a year to answer "who has laptop DELL-4471" in 2026.

Now meet Snipe-IT.

A free and open source IT asset management platform that runs on your own server and tracks every laptop, phone, license, and cable your company owns. No cloud. No agents. No per-seat fee.

Built in 2013 by a developer named Alison Gianotto who was the CTO of a New York City ad agency when the office had to move. Her team was tracking hundreds of laptops in a Google Doc. She tried every commercial and open source tool. None of them worked. So she wrote her own.

When she finished the first full inventory, she discovered the agency was missing $25,000 in assets.

Her husband nagged her for months to add a PayPal button. She finally did it. Two days later she had a paying customer. Snipe-IT itself stayed free.

14,113 stars. 3,886 forks. AGPL-3.0. Pushed to the repo today.

Here is what Snipe-IT gives you:

- Track every laptop, phone, monitor, cable, and accessory in one place
- Assign assets to users, departments, or locations with full history
- Software license management with seat tracking and expiration alerts
- Barcode and QR code generation for physical labels
- Custom fields, statuses, and categories
- LDAP, SAML, and Google Workspace SSO
- REST API for every endpoint
- Runs on any Linux box, Docker, or a $5 VPS

Here is what Snipe-IT costs:

Zero. Forever. No per-user fee. No asset cap.

ServiceNow charges $100 per user per month. Snipe-IT doesn't.
Lansweeper charges $239 a month for 2,000 assets. Snipe-IT doesn't.
Freshservice meters every asset over 500. Snipe-IT doesn't.
Ivanti and Flexera hide their pricing behind a sales call. Snipe-IT doesn't.

Here's the wildest part:

Snipe-IT is used by Sony, Yale University, Fujitsu, and thousands of school districts, hospitals, and government agencies that couldn't stomach a $124,000 ServiceNow bill. Alison's official title at Grokability is Chief Mohawk Officer.

For a 50-person IT team, you save $60,000 a year.
For a 500-person company, you save $600,000 a year.

Your assets. Your database. Your server.

100% Open Source. (Link in the comments)
图片
9 30 10.6K
Akshay 🚀 AI @akshay_pachaar · 2026年07月28日 12:10
一个刁钻的 LLM 面试题:

你在 vLLM 上部署推理模型,遇到长链推理老是爆显存。

于是你加了 KV cache 压缩,删掉了 90% 的缓存 token。

结果显存占用纹丝不动,GPU 还是爆了。

这是为什么?

(答案在下面)

删掉 90% 的 KV cache 可能几乎释放不了任何显存。

这听起来很反直觉,但这是由目前生产级服务器存储缓存的机制决定的。

KV cache 随 token 生成而增长。每个 token 都会在每一层追加 key 和 value 向量,生成过程中没法释放。

这是推理模型最主要的内存成本。

如果一个 32K token 的 CoT 缓存了约 32K token 的 KV 向量,那么 4-bit 量化的 Qwen3-32B 在 24GB 的 GPU 上,到 24K token 左右就会爆显存。

一个显而易见的方案是:只保留重要 token,扔掉剩下的,毕竟注意力机制足够稀疏,允许这么做。

但即便如此,也还没解决显存问题。

原因在于 Paged Attention,它是 vLLM 等大多数生产服务器背后的内存管理器。

在底层,它把 GPU 显存切成固定的物理块(Block),每个块存大约 16 个 token 的 KV。

只有当一个块内的所有位置(Slot)都变空时,这个块才会归还给分配器。

因为剔除逻辑是按“重要性”挑选的,而这些重要 token 散落在各个块里……

……所以即便剔除了大部分,几乎每个块里都还残留着一些“幸存”的 token。

比如,你在 1,000 个块里删了 16,000 个 token 中的 14,000 个,大概率每个块里还是剩了至少一个 token。

这意味着分配器几乎释放不了任何空间。

把新 token 塞进这些空出来的槽位也不理想,因为会破坏缓存的布局。

假设第 16,001 个 token 来了,把它放在原先第 40 个 token 的位置。缓存读取顺序就变成了 38, 16001, 41…… 顺序乱了。

注意力机制确实还能算对,但前提是每个槽位都得额外记录它到底对应哪个位置。

这引入了额外的簿记开销,而顺序布局本身是可以避免这些的。

所以,逻辑上缓存小了 90%,物理上还是那么大。很多关于压缩的研究忽略了这一点,因为他们是在预分配的连续张量上测的,而不是在分页服务器上。

此外还有个问题。

剔除方法通常根据注意力得分(Attention Scores)来选 token。

但生产环境中常用的 FlashAttention 等快速算子,根本不保存这些得分。

它们分片计算注意力,随算随扔,这也是它们快的原因。

所以,剔除所需的精确信号在显存里压根拿不到。折中的办法是退回到原生注意力(Eager Attention)去构建完整矩阵,但这又牺牲了 FlashAttention 带来的速度。

NVIDIA 发布了一个叫 TriAttention 的方法来解决这两个问题。

它不需要注意力得分。相反,它在 RoPE 应用之前,根据 key 和 query 向量的几何结构来打分,因为这些向量在稳定的聚类中。

针对内存问题,它每解码 128 个 token 运行一次“整理压实”(compaction pass)。

幸存的 token 会向前滑动,填补剔除留下的空洞,这样整块的显存就能清空并还给分配器,同时保持缓存的顺序。

在长推理任务中,该方案在保持全注意力精度的同时,解码速度提升了 2.5 倍,KV 显存占用减少了 10.7 倍。

KV cache 压缩是一个巨大的基建问题。决定它是否有效的数字是“释放了多少块”,而不是“删了多少 token”。

你可以在这里找到 NVIDIA 的详细文章:https://research.nvidia.com/labs/eai/blogs/kv-cache-compression-and-its-infra-problems/

我写了一份关于 KV cache 工作原理的第一性原理拆解。涵盖了模型为什么要存 KV、缓存为什么随 token 增长,以及有无 KV cache 的生成速度对比。

点击下方阅读。
A tricky LLM interview question:

You're serving a reasoning model on vLLM, and it keeps running out of GPU memory on long traces.

So you add KV cache compression and evict 90% of the cached tokens.

VRAM usage stays as is and GPU still runs out of memory.

Why?

(answer below)

Evicting 90% of the KV cache can free almost none of the memory it was using.

This sounds counterintuitive, but it follows directly from how production servers store the cache today.

The KV cache grows with every token a model generates. Each token appends its key and value vectors across every layer, and nothing is freed while generation continues.

This is the dominant memory cost for reasoning models.

If a 32K-token CoT caches ~32K tokens of KV vectors, a Qwen3-32B with 4-bit weights will run out-of-memory around 24K tokens on a 24GB GPU.

One obvious solution is to keep the important tokens and drop the rest, since attention is sparse enough to allow it.

But this does not solve the memory problem yet.

The reason is paged attention, which is the memory manager behind vLLM and most production servers.

Under the hood, it splits GPU memory into fixed physical blocks, each one holds the KV for about 16 tokens.

This block returns to the allocator only when every slot inside it is empty.

Since the eviction logic selects tokens by importance, and such tokens are scattered across blocks...

...so despite eviction, almost every block is left with at least some survivor tokens.

For instance, if the logic evicts 14k of 16k tokens across 1,000 blocks, most likely every block will still have a token.

This means the allocator frees almost nothing.

Placing the new tokens into those freed slots is not ideal because it breaks the cache's layout.

Say token 16,001 arrives, and it's placed in the slot the 40th token used to hold. The cache now reads position 38, then 16,001, then 41, so the cache is no longer in token order.

Attention can still compute the right answer from that, but only if every slot now carries a separate note recording which position it actually holds.

This introduces another bookkeeping cost that an in-order layout inherently avoids.

So the cache is logically 90% smaller and still physically the same size. Many compression results miss this because they measure on pre-allocated contiguous tensors rather than a paged server.

There's another problem.

Eviction methods pick which tokens to keep by looking at the attention scores themselves (as expected).

But fast attention kernels used in production, like FlashAttention, never save those scores.

They compute attention in small pieces and throw the full score grid away as they go, which is also why they're fast.

So the exact signal eviction methods need isn't available in memory. The workaround is to fall back to eager attention and build the full matrix, which gives up the speed FlashAttention was there to provide.

NVIDIA published a method called TriAttention to solve both these problems.

It never needs attention scores. Instead, it scores tokens from the geometry of the model's key and query vectors before RoPE is applied, where those vectors sit in stable clusters.

For the memory problem, it runs a compaction pass every 128 decoded tokens.

The surviving tokens slide forward to close the holes eviction creates, so whole blocks empty out and return to the allocator while the cache stays in token order.

On long reasoning traces, the approach matches full-attention accuracy while decoding 2.5x faster and using 10.7x less KV memory.

KV cache compression is a big infrastructure problem. The number that decides whether it works is the count of freed blocks, not the count of evicted tokens.

You can find the NVIDIA write-up here: https://research.nvidia.com/labs/eai/blogs/kv-cache-compression-and-its-infra-problems/

I wrote a first-principles breakdown of how the KV cache works. It walks through why the model stores keys and values at all, why the cache grows with every token, and a comparison of LLM generation speed with and without KV caching.

Read it below.
9 19 19.1K
Huan AI @Huanusa · 2026年07月28日 04:31
安徽农民在日本打工,工头小林君看他工资低,经常给他多算加班费。
他困惑的表示:
也不知咋回事,电视里看到的鬼子在现实里一个也没看到。 https://x.com/Huanusa/status/2081960589694468412
19 6 18.9K
余温 AI @gkxspace · 2026年07月28日 03:03
我用一句话,让 Codex+Higgsfield 上线了一款能联机玩的武侠游戏!!!

全球联机,进同一片水墨竹林,轻功跳屋顶、剑气波隔空对轰、连斩有称号播报,手机也能玩。

整个过程我没画一张图、没写一个音符、没租一台服务器:
1、Nano Banana Pro 出水墨剑客概念图
2、Image to 3D 把图直接变成带骨骼的 3D 模型
3、3D Rigging 从动作库套上出剑、轻功、疾风步
4、Sonilo 生成古琴 BGM,Mirelo 出剑气音效
5、代码让 Cursor 里的 Fable 5 写,它自己调 Higgsfield CLI 把这些全串起来
6、最后 higgsfield game deploy 一条命令上线,域名和全球联机服务器全自动,再一条 publish 直接上架 marketplace

以前这件事需要找游戏工作室,美术、作曲、3D建模、动画、全栈开发一个都不能少。

但现在真是门槛无限低啊,想玩什么游戏直接自己做~
22 19 10.5K
阿西_出海 AI @axichuhai · 2026年07月28日 02:37
发现一个适合一人公司的开源神器—MetaGPT,在 GitHub已斩获 6.9w+star

你只需一句话提需求,它就会自动拆解任务

产品经理完善需求、架构师设计结构、项目经理统筹安排、工程师负责出代码、QA给出测试用例,整个研发团队的活全包了

过去需要一个团队花几天甚至几周的开发工作,现在一个人就能快速搞定
15 22 10.3K
Nikita Bier AI @nikitabier · 2026年07月28日 16:24
上一代人通过照片和视频表达自我。

而现在的表达方式,是软件。

谁将创造下一个直接在 X 时间轴里运行的现象级游戏或应用?

快来试试全新的 Grok 应用构建器:https://x.com/grok/status/2082134072793637196
Photos & video were how the last generation expressed themselves.

Today’s form of expression is software.

Who will create the next viral game or app that plays right inside of the X Timeline?

Try the new Grok app builder. https://x.com/grok/status/2082134072793637196
458 136 55.9K
火山哥🕊️ 加密 @huoshan007 · 2026年07月28日 08:28
我表弟高中还没毕业,已经把我看沉默了。

他自己捏了个美女图,丢给 Claude 做成 AI 角色,再顺手生成跳舞视频,发 TikTok。

18 天,涨了 61267 粉。

然后开 Fanvue,9.9 美元订阅,768 个人付费。

半个月到手差不多 7603U。

最离谱的是,他还觉得自己只是随便玩玩。

我真的服了。

AI 网红这条线,很多人还在研究,他已经开始收钱了。
42 11 22.6K
0 时间范围 点赞 > 分类 排序
历史账号
默认展示全部账号;可展开上方账号区多选,或用「分类」下拉按类筛选,选好后点右上「更新所选」开始拉取历史爆文。
(本页面独立,需手动更新才会爬取,不浪费 API)
⚙️ 设置
更新于 2026-07-28 20:30 UTC
📡 监控账号
🕐 抓取参数
小时
小时
💰 价格展示代币
🐊 Meme大神
扫描SOL/BSC/Robinhood热门代币,自动发现金狗钱包,入库到爆文列表「其他」分类
加载状态中...
小时
U
U
人民币(严格超过)
个钱包
条/链
📱 自媒体 API 设置
Threads
币安广场
微信公众号
推特数据 API(twitterapi.io,会员默认数据源)
爬取推文/账号的主力数据源 · 注册:twitterapi.io ↗
🤖 AI 大模型 API(改写/翻译,会员默认 uau.cc)
用于推文 AI 改写功能(兼容 OpenAI 接口格式)· 注册:uau.cc ↗
夸克网盘 Cookie(用于转存)
用于一键转存推文中的夸克分享链接到你的网盘
✍️ AI改写提示词
🔐 修改账号密码