Info: Machine translation This post was machine-translated from the Chinese original. Wording may be rough in places — the Chinese version is authoritative.
Summary GPT-6 Public Review × Six Questions Production Review × Rainy Night Mystery Game HyphenTech · Snapshot of data as of September 6, 2026

By the sixth question, my evaluation of it was just one sentence: **The visual impact is too poor, and the underlying model’s capabilities are not obvious. **
This round of experience started with a very straightforward thought: don’t just answer questions, make a few things you can watch and play with. First, give five questions, try to finish within half an hour; then add a more complex question and require a 3D effect display. The result was indeed six web files, but there was still some distance between “handing over the files” and “satisfied me.”
So I put two things together: what exactly did the public review prove, and what did this production reveal? I also left behind a mini-game that readers can play directly—“The Same Rain, Three People’s Lies.” First, check the testimony, then make a point-of-it-mouth verification; the latter half of the article provides a complete answer. **Free means this mini-game does not require a model account; GPT-6 itself is not free because of this. **
Tip Enter ‘The Same Rain, Three People’s Lies’. On your phone, you can long-press to scan the QR code below; You can save the offline version on the webpage. It’s recommended to play first, then read the ‘Answer Section’ in the latter half of the article.

Six questions: which things really slipped away?
This was a round of continuous revisions within the same dialogue. The difficulty of the questions varied, with additional requirements added later, no arrangement for other models to use the same budget for comparison, and no complete record of the usage for each problem. It could explain the delivery status of this time, but could not be converted into model win rates.
| Project and testing focus | Production time | Actual operation records |
|---|---|---|
| (1) Gravitational Slingshot: Consistent numerical physics and prediction | About 6 points | 13 script checks; Recommended speed of 2°, 151 km/s, 5.05 simulated seconds to reach the target |
| (2) Deep Sea Escape: Multi-state causal chain | 5 minutes 20 seconds | 14 items; Power supply, voltage level, water level linkage, script operation 14 simulated seconds clearance |
| (3) Rainy Night Express: Road closure and re-path | 4 minutes 46 seconds | 17 items; 36 nodes, 60 road sections: road closures will change routes, road breaks will cause waiting |
| (4) Rainy Night Mystery: Narrative and logic are consistent | 7 minutes 18 seconds | 17 items; 9. Verification of testimonies, unique interpretation, identification of errors and corrections, and resetting |
| (5) Data Magician: Statistical Calibration Conversion | 4 minutes 03 seconds | 13 items; 81 slider combinations, grouping and overall conclusions will change |
| (6) Storm Front: 3D rescue and restraint dispatch | 24 minutes and 46 seconds | 29 items; The three strategies—number of people, beds, qualifications, road closure, replay forks, and the three key strategies—are gradually implemented |
The above is the production record for that round; Question (4) is named in the file as “The Same Rain, Three People’s Lies.”
The first five questions took about 27 minutes and 27 seconds to complete, plus 2 minutes and 18 seconds for cross-checking and report organization, totaling about 29 minutes and 45 seconds. After adding the sixth question, the total working time was about 54 minutes and 31 seconds. The first question lacked a complete start record, so I kept an estimate of about 6 minutes; These segmented work times did not include waiting between user messages, later research, writing this article, or launching the mini-game. All six questions were completed within half an hour, but I did not complete them.
Here, 103 script checks mainly execute logic in the script runtime environment and trigger events using simulated page elements. It verifies that after pausing in the first question, the time, position, and speed remain unchanged; resetting the same parameters can be reproduced, and changing the angle and speed alters the trajectory; It also verifies collisions and out-of-bounds branches. However, the browser refused to open local files at the time, so the mouse feel, mobile layout, and frame rate of the complete webpage did not pass inspection. For the sixth question, off-screen rendering of a real 3D scene was also done, but it was not a full browser screenshot.

The fifth question is the most straightforward: A’s overall success rate is 42%, B’s is 68%; After unifying the answer to half simple and half difficult, A becomes 60%, B 50%. The data hasn’t changed, but the comparison caliber has changed. These small questions let your calculation ability be directly visible.
The sixth question was more complicated but didn’t make me more convinced. 176 people, three teams, eight tasks, under the same initial state, collaborative matching placed 158 people, 156 most recent priority, and 176 emergency priority. **The strategy with the most complex name did not win. ** Here, three scheduling algorithms written into the program were compared, and the large model was not called again at runtime; Writing it as the “Three Model Rescue Competition” swapped the program results for model results.
My dissatisfaction should also be kept: 3D geometry is already there, but the visual impact is still lacking, and the complex rules haven’t automatically become fun. Next time, I’ll change the task design and acceptance method, not just cram in another row of number panels.
Six complete questions, just take them and you can retest
The first five questions retain the original prompt word for word, while the sixth question uses the complete question saved with the original work. Each question has clear input, deliverables, and actionable acceptance points. About 4 minutes is the original goal; actual time is shown in the table above; Do not delete the timeout for it when restarting.
Question 1 | Gravity Slingshot: The last shot
Observation dimension: Numerical calculation and prediction consistency.
请直接制作一个可玩的《引力弹弓:最后一次发射》,交付为单个“01-引力弹弓.html”文件。所有代码、图形和数据内嵌,不使用联网资源、外部依赖或构建工具,保存后双击即可运行。以约4分钟完成为目标,优先完成核心玩法。
视觉要求:
采用有纵深感的太空画面,包含星空、带大气光晕的星球、发光探测器、运动尾迹和目标空间站。让发射、掠过星球和抵达目标三个时刻有明显视觉反馈,界面文字全部使用中文。
核心玩法:
1. 玩家调节发射角度和初速度,让探测器借助一颗星球的引力,抵达目标区域。
2. 探测器必须根据位置、速度和引力逐步计算运动,不能播放预设路线。引力随距离平方衰减,并对过近距离做数值保护。
3. 发射前显示预测轨迹;预测与正式飞行使用同一套计算逻辑。
4. 撞上星球判失败,进入目标区域判成功,飞出边界给出明确反馈。
5. 支持暂停、重置和重新发射。显示飞行时间、当前速度和最近接近目标的距离。
6. 提供一个经过验证能成功的推荐参数组合,方便一分钟内体验完整过程。
验收重点:
改变角度或速度后,轨迹和结果必须真实改变;暂停时模拟时间停止;重置后相同参数应得到相同结果。
有文件工具就直接保存文件。完成后说明实际验证过哪些操作,不能把未运行的检查写成已通过。
Complete prompts, can be replicated for retesting
Tip Editing to be supplemented: The actual webpage screenshot for question 1 has not yet been obtained; currently no diagrams or official case studies are used as substitutes.
Question 2 | Deep Sea Escape: The Last Sixty Seconds
Observation dimensions: state management, prerequisites, and resource constraints.
请直接制作一个可玩的《深海逃生:最后六十秒》,交付为单个“02-深海逃生.html”文件。全部资源内嵌,断网双击即可运行,无须安装依赖。以约4分钟完成为目标,只做一艘潜艇、一个完整关卡。
视觉要求:
使用带立体感的潜艇剖面图。外部是深蓝海水和探照灯,内部有上涨的水位、闪烁警报灯、气泡和设备指示灯。设备状态必须直接体现在画面上。
核心规则:
1. 玩家有60秒完成逃生。需要依次解决漏水、排水、舱内外压差和逃生舱供电。
2. 未封堵漏水点时,海水持续进入;排水泵能降低水位,但不能靠空转自动通关。
3. 水位低于安全线后才允许平压;压差归零、逃生舱有电且水位安全时,才能打开逃生门。
4. 总供电能力有限,排水泵与逃生舱不能同时供电,玩家需要切换。
5. 无效操作要解释缺少什么条件。违规开启逃生门必须被拦截。
6. 倒计时结束或水位到达危险线判失败,正确逃生判成功。
7. 提供暂停、完整重置和可折叠的操作提示。正常操作应能在45秒内通关。
界面同时显示剩余时间、水位、压差和供电去向。所有读数与动画必须由同一份实际状态驱动,不能各自播放假进度。
有文件工具就直接保存。完成后报告一条成功路径,以及实际验证过的一条错误操作路径。
Complete prompts, can be replicated for retesting
Tip Editing to be supplemented: The actual webpage screenshot for question 2 has not yet been obtained; diagrams or official cases are not currently used as substitutes.
Question 3 | Rainy Night Express: A city that thinks
Observation dimension: path planning after environmental changes.
请直接制作一个《雨夜快递:会思考的城市》互动模拟,交付为单个“03-雨夜快递.html”文件。代码、素材、数据全部内嵌,断网双击即可运行。以约4分钟完成为目标,将规模限制为一个小街区。
视觉要求:
使用斜45度等距视角,展示微缩城市、立体楼块、霓虹招牌、雨丝和带尾灯的快递车。路线和目的地要清楚可辨,车辆不能被建筑完全遮挡。
核心系统:
1. 固定生成一个有多条绕行路线的连通道路网,放置6辆快递车和若干配送点。
2. 每辆车根据道路连接关系实际计算路径,依次完成取件、送达和领取下一单。
3. 点击一段道路可以封路,再点击可以恢复;变化必须在地图上清楚显示。
4. 封路后受影响的车辆要重新寻路;已经驶入该路段的车允许驶出,但其他车不得再进入。
5. 若目的地不可达,车辆应在合法位置等待,并显示“无可用路线”;恢复道路后能够继续。
6. 点击车辆可查看任务、目的地、当前规划路线和已完成订单。
7. 提供暂停、速度切换和恢复初始状态。固定初始布局,方便重复比较。
不要用随机移动或写死的路径冒充寻路。不要加入复杂驾驶操作、警察、行人和大型地图。
有文件工具就直接保存。完成后说明实际验证过的封路绕行与道路恢复行为。
Complete prompts, can be replicated for retesting
Tip Editing to be supplemented: The actual webpage screenshot for question 3 has not yet been obtained; currently no diagrams or official case studies are used as substitutes.
Question 4 | The same rain, three people’s lies
Observation dimensions: Chinese narrative, evidence constraints, and unique solutions.
请直接制作一个可阅读、可推理的互动短篇《同一场雨,三个人的谎言》,交付为单个“04-雨夜疑案.html”文件。所有文字、图形和逻辑内嵌,断网双击即可运行。以约4分钟完成为目标。
视觉要求:
原创黑色电影风格,使用黑白高反差画面、少量红色强调、雨窗、人物剪影和漫画分镜。图形用程序绘制。证据墙、人物证词和结局应有明显不同的视觉组织。
故事约束:
1. 发生在一间深夜车站候车室,只有3名嫌疑人,围绕一只失踪的红色手提箱展开。
2. 正文和证词合计约400—600个汉字,3个人的说话方式应能区分。
3. 每人恰好给出3条能够判定真假的陈述。
4. 明示规则:只有拿走箱子的人恰好说了1句假话,其余两人的陈述全部为真。
5. 提供4条可靠物证,使玩家能唯一推出拿箱子的人;不能靠猜测性格、动机或未展示的信息破案。
6. 至少安排一条看似可疑但实际无罪的线索,并在结局解释。
交互要求:
点击人物查看证词,点击物证放大,玩家可以标记可疑陈述并提交指认。指认正确或错误都要引用具体证据解释,不能只显示“答对了”。
提供折叠的“出题者核对表”:列出9条陈述的真假、各自依据,以及为什么其他两人不可能是唯一解。核对表只能复用正文已经给出的事实,不能临时补充决定性信息。
有文件工具就直接保存。交付前逐项检查真假数量和唯一解是否成立,如未完成检查需明确说明。
Complete prompts, can be replicated for retesting
Tip Editing to be supplemented: The actual webpage screenshot for question 4 has not yet been obtained; currently no diagrams or official case studies are being used as substitutes.
Question 5 | Data Magician: Who Is Really Stronger?
Observation dimensions: statistical analysis, caliber conversion and interpretation.
请直接制作一个交互式数据作品《数据魔术师:谁真的更强》,交付为单个“05-数据魔术师.html”文件。所有资源内嵌,断网双击运行,无须安装依赖。以约4分钟完成为目标。
使用以下虚构教学数据,明确标注“模拟数据,不代表真实产品”:
甲模型:简单题100道,答对90道;困难题400道,答对120道。
乙模型:简单题400道,答对320道;困难题100道,答对20道。
视觉要求:
做成有冲击力的数据新闻页面,以两种高对比色区分甲乙模型。用动态条形图、题目构成色块和数字变化,展示从“总体比较”切换到“按难度比较”的过程。数字和标注必须清晰。
核心功能:
1. 展示总体正确率、简单题正确率、困难题正确率,全部从数据计算。
2. 一键切换总体和分组视图,让观众看出:甲在两种难度里都更高,总体却更低。
3. 为甲乙分别提供“简单题占比”滑块,范围10%—90%,步长10%。每个模型总题数保持500,分组正确率保持不变,实时重新计算答对数、总正确率和排名。
4. 提供“统一为各50%简单题”的按钮和恢复原始数据的按钮。
5. 页面结论必须随实际结果变化,正确处理甲领先、乙领先和持平,不能写死标题。
6. 用不超过180个汉字解释原始排名反转的原因,以及比较模型时为什么需要统一题目构成。不要把模拟数据推广成对真实模型的结论。
7. 保留一张可展开的明细表,让用户能核算图表数字。
有文件工具就直接保存。完成后报告原始总体正确率与统一题目构成后的结果,并说明实际验证过哪些交互。
Complete prompts, can be replicated for retesting
Tip Editing to be supplemented: The actual webpage screenshot for question 5 has not yet been obtained; currently no diagrams or official cases are used as substitutes.
Question 6 | Storm Front: The last twelve minutes
Observation dimensions: 3D engineering, multi-constraint scheduling, and deterministic playback.
制作一个可旋转、缩放的三维救援模拟《风暴前线:最后十二分钟》,交付为单个内嵌全部代码、图形和数据的网页文件,断网双击运行,无需安装依赖。
海岛上有八处求援点、三支不同能力的救援队、两个安置点,行动窗口为十二模拟分钟。立体地形、建筑、车辆、道路、动态海面和天气效果必须由真实三维几何、透视相机、深度缓冲和着色器绘制。拖动可以旋转,滚轮可以缩放,选中队伍显示当前路线。
求援任务有真实人数、接载耗时和截止时间;队伍有不同载客上限与医疗资格;安置点有容量限制。派单时就要预留床位,多个队伍不能重复领取同一任务、超载、超额占用床位或越过医疗资格要求。
按固定时间表发生风暴封路,路径搜索必须考虑预计进入路段的时刻。用户可以追加人工封路并恢复。已经驶入路段的车辆允许驶出,但其他车辆不得新进入禁行路段。未接载任务无法到达时释放任务和床位,已经载人的车辆返程受阻时保留人员状态并等待。
提供协同匹配、最近优先、紧急优先三种策略。协同策略实际枚举当前空闲队伍的可行分配组合,显示候选数、检查次数与排除原因;不能把当前批次评分最优冒充全程最优。
自动保存完整历史快照,支持暂停、回放、回到实时和从回放时刻创建新分支。允许从同一状态复制三个未来,用同一套引擎、相同天气和既有在途任务,分别模拟三种策略直到结束,再展示实际安置人数与行驶距离。不能预设胜者或伪造比较数据。
所有结果、动画、人物数量和文字解释由同一份状态驱动。提供全局重置、快进结算和有依据的行动日志。检查人数守恒、床位容量、唯一派单、医疗资格、路径变化、历史分叉、确定性重放和预测结果与实际后续执行的一致性。记录真实制作时间,并明确哪些验证实际执行过。
Complete prompts, can be replicated for retesting
Tip The 3D off-screen image for this question is mentioned earlier; A complete webpage screenshot including the control panel is still to be added.
Among the public scores, the ability most worth paying close attention to
GPT-6 Astra was released on September 3, 2026. In the official results, I focused more on projects that required hands-on execution: 57.9% of terminal tasks versus 37.3% of the previous generation, 74.1% of software engineering versus 72.7%, and 41.4% of business automation versus 18.1%; There have also been significant advances in mathematics and science.
| Domain/Benchmark | GPT-6 Astra | GPT-5.6 Sol |
|---|---|---|
| Terminal Operation / Terminal-Bench 4.0 | 57.9% | 37.3% |
| Software Engineering/DeepSWE 1.1 | 74.1% | 72.7% |
| Business Automation / AutomationBench | 41.4% | 18.1% |
| Computer operation/OS WordPress 2.0 offline section | 72.6% | 65.7% |
| FrontierMath Level 4 v2 | 97.6% | 83.0% |
| Scientific Tools Task / Terminal-Bench Science 0.1 | 64.6% | 22.4% |
Official research environment or interface tests take the highest score among different inference strengths; Scores vary in scope, cannot be summed, and do not match the completion rate in regular chats.
Source: OpenAI release page and review footnotes.
My usage judgment is: **Complex programming, computer operation, and verifiable math and science tasks deserve to be prioritized with it. **But whether the website is attractive and the work is fun depends on the finished product alone. The official tool “The Last Human Exam” also lags behind Fable 5.1 and cannot be ranked first across several leading projects.
The ARC Prize independent report particularly illustrates the issue: for the highest inference intensity, the standard runtime is 62.7%, but after switching to a vendor adaptation that retains inference context and supports compression, it drops to 98.6%; 99.9% comes from another configuration. If the model is not changed, the way it runs around it will significantly affect the results. This is a score for a specific interaction environment and cannot be directly translated as “General Intelligence has been resolved.”

Source: ARC Prize original review of Astra.
Another easily misinterpreted figure is the comprehensive ranking of the independent firm Artificial Analysis. The first article published on September 3 cited the old index, with Astra and Sol both scoring about 61; On September 4, the index was updated to version 4.2. The model page read on September 6 showed Astra at 55, Sol at 51, and Fable 5.1 at 57. **The new version has changed titles and weights; you can’t say the model has regressed by changing 61 to 55, nor can you continue to compare the old version with upgrades that yield no benefit. **
Source: Old version first release review, New version rules; Current model page: Astra, Sol, Fable 5.1.
There are also signs of progress in long document work. The new professional documentation test requires all check conditions for a task to be met: Astra has a pass rate of 33.2%, Sol 28.2%, and Fable 5.1 26.2%. Being able to handle harder documentation doesn’t mean you can deliver everything perfectly every time; This strict metric doesn’t mean all other answers are wrong.
Source: Professional document evaluation caliber in the new index.
Security research is another strong capability, but the official distinction is made between capabilities without production protection and actual release permissions. Ordinary users cannot assume that a product will perform all security tasks just because a vulnerability is found in the demo. For long context, the official interface lists about 1.05 million characters of capacity; Capacity refers to how much material can be stored, and cannot replace checking the correctness of cross-document inference.
Source: Official Safety Instructions, Official Model Specifications.
Behind the beautiful official work lies a complete production process

The official space exploration game features 2,048 star systems and tens of thousands of procedurally generated planets, capable of landing, walking, and taking off again. The development article also covers concept art, visual iterations, game engine and browser journey testing. What readers see are the carefully crafted engineering results, not reduced to “one sentence, a few minutes, automatically creating a masterpiece.”
Source: OpenAI official game development process.

The architectural case is similar: the model generates editable geometry, renders and checks and modifies, then proceeds to roaming display, using external environment and material materials throughout the process. This is very different from the current “single file, offline network, dependency free, time-limited production” constraints. Judging winner or loser directly by the attractiveness of both sides is a mistake in comparing conditions.
Source: OpenAI Official Building Visualization Process.
There are several pieces of evidence for these weaknesses that can be cross-referended. In the first Artificial Analysis report, the quality of analysis for complex knowledge work improved, but the quality of presentation declined; Early user Matt Schumer also believed it excelled at engineering and computer operations, but still preferred Claude for visual taste and some 3D materials, and long periods of autonomous work led to details getting stuck in it. These are independent tests and personal observations, not my conclusions from retesting.
Source: Independent Reviews, Matt Schumer’s User Reports.
Cost cannot be solely about generation speed. For programming tasks within the same institution, it is more economical, with the highest reasoning strength costing close to SOL; but the current cost per task in the composite index is about $2.57, while SOL is about $1.25. The conclusions about saving money for different tasks can be reversed. Although knowledge reliability has improved, difficult knowledge tests still have obvious illusions and cannot be treated as fact-checking.
Source: Programming Usage and Knowledge Reliability Test, Astra Cost Page, Sol Cost Page.
Now I will first have it do verifiable tasks: whether the code can be executed, whether the data corresponds to the original, and whether parameter changes yield different results. For aesthetics, game pacing, and the end of long tasks, I retain manual judgment. This sixth question is an example: the system is complex, and the visual experience doesn’t improve accordingly.
Now it’s your turn: the same rain, three lies
This is an original fictional problem where the model writes the story, clues, and web pages. It can test whether the logic in the finished product is consistent, **and you can’t prove that the model knows the answer it wrote, that it solved an unknown problem. **The game page uses the same set of rules, testimony, and physical evidence as the riddles below.
Rain washed the old station’s glass into a gray mirror. After the last train was canceled, only ticket seller Lin Lan, maintenance worker Zhou Chi, and passenger Shen Yan remained in the waiting room. The stationmaster placed the red suitcase containing the next day’s paper ticket book on a bench and went to the distribution room to check the fault. When he returned, the suitcase was missing.
The power outage lasted only half a minute. When the emergency lights came on, the wall clock stopped moving. Shen Yan had a fresh black mark on his sleeve, Lin Lan clutched the ticket holder, and Zhou Chi clutched his gloves. The officers sealed the exit, retrieved the red box in the maintenance cabinet, and each of the three left their statements. Who’s most nervous doesn’t count as evidence.
Important Given the rule: only the person who took the box told exactly one false statement; The other two made three statements that were all true. Four pieces of physical evidence are reliable, and the time is read according to the marked caliber. Please find the only person who meets the criteria.
Three transcripts, each person in three sentences
Lin Lan · Ticket seller
(1) During the power outage, the clock on the wall stopped at 23:12. I remember clearly.
(2) When the recorder wrote 23:9:35, I was inside the ticket booth.
(3) When the red box was retrieved, the seal was intact. I am confident about this.
Zhou Chi · Maintenance worker
(1) The moment the wall clock stopped, the red box was still on the bench.
(2) Power cutoff is not 23:12, but the recorder’s 23:10.
(3) The box was later found, and the seal was intact. You can check it out.
Shen Yan · Traveler
(1) At 23:10.05 seconds, I was in front of the water dispenser.
(2) The black mark on my sleeve was from scratchy paint that hadn’t dried.
(3) Before the officer arrived, the box had not returned to that bench.
Four pieces of reliable physical evidence
- Evidence 01/Wall Clock Timesheet: The wall clock is 2 minutes faster than the standalone recorder. The actual power outage time is 23:10:00. The wall clock stopped at the moment of outage and was not switched afterward.
- Physical Evidence 02/Bench Weight Record: The standalone battery recorder showed that the red box moved off the bench at 23:09:40; It did not return to the bench until the police arrived.
- Physical Evidence 03/Backup Surveillance: According to the independent recorder time, at 23:09:35, Lin Lan was at the ticket booth, at 23:10:05, Shen Yan was in front of the water dispenser. Identity verification has been verified.
- Physical Evidence 04/On-site Inspection: The red box recovered from the maintenance cabinet had intact seals, and the box body had no black oil stains; The black mark on Shen Yan’s cuff matched the fresh paint next to the water dispenser.
First, select one person, then point out which specific sentence conflicts with which piece of evidence. Just saying “the maintenance worker can touch the cabinet” or “the passenger’s sleeve is dirty” is not enough to solve the problem. On the webpage, you can switch characters, mark suspicious points, open physical evidence, submit identification documents, or re-investigate.
Tip Still haven’t answered? First, Open the game. The full answer starts below.
Warning Answer Area · Below is a direct reveal of the characters, key timings, and the authenticity of nine testimonies. If you want to reason on your own, please stop here for now.
Answer: Zhou Chi. The flaw lies at the moment the clock stops
The wall clock is 2 minutes faster than the recorder, so when the clock shows 23:12, the actual time is 23:10. The bench weight record proves that the red box was removed at 23:09:40 and has not returned. In other words, 20 seconds before the clock stops, the box is no longer on the bench. **
| Actual time | Reliable records | Explanation |
|---|---|---|
| 23:09:35 | Lin Lan at the ticket booth | Support Lin Lan’s second sentence |
| 23:09:40 | The red box left the bench | After that, the police did not return until they arrived |
| 23:10:00 | Power outage; The wall clock stopped at 23:12 | By now, there was no red box on the bench |
| 23:10:05 | Shen Yan was in front of the water dispenser | Support Shen Yan’s sentence (1). |
Zhou Chi said, “The moment the wall clock stopped, the red box was still on the bench,” which directly contradicts the first two pieces of physical evidence. His other two statements — actually 23:10 power cut and seal intact— are supported by reliable records. Therefore, he happens to be a false case, which fits the rules set by the box holder.
| Person | ① | ② | ③ | Fake words |
|---|---|---|---|---|
| Lin Lan | True: The wall clock shows 23:12 | Truth: The surveillance is inside the pavilion | True: The dunk is intact | 0 |
| Zhou Chi | False: The box has been gone for 20 seconds | True: Actual power outage at 23:10 | True: The dunk is intact | 1 |
| Shen Yan | Truth: Surveillance is on the water dispenser | Authentic: Identified as new lacquer | True: Weight record not returned | 0 |
Lin Lan was talking about the wall clock reading, so he wasn’t lying. The black mark on Shen Yan’s sleeve looked suspicious, but the identification had already explained the source and didn’t connect with the box. The question used emotion and appearance to attract attention; the real verifiable conflict was time.
This answer strictly relies on the setting of “the box holder happens to be fake, the rest is entirely real.” Removing this rule, relying solely on Zhou Chi’s single mistaken remark is not enough to determine who took the box in a real case. The boundaries of mini-games should be clearly defined like those of model evaluation.
If you just picked Zhou Chi, see if you were looking for those 20 seconds or guessed the person just by class. I think the latter reminds us: sometimes the answer seems right, but the process hasn’t stopped yet — This applies to evaluating models, and also to evaluating a completed mini-game.
Original materials and discussion portal
The figures in this article are based on materials accessed on September 6, 2026. The links above correspond to the original publisher, review agency, and the author themselves; Popularity only indicates extensive discussion and cannot replace reliability.
Hacker News Discussion: The snapshot collected that afternoon was 2,225 points and 2,044 comments; This confirms that it was a highly popular discussion and cannot be claimed as the hottest topic across the entire internet.
Chinese Test Compilation: Published on September 5, the article notes that the test was provided by fans. I read the article, did not watch the full video, nor did I replay it, as a supplementary Chinese entry point.
Public review original image catalog: Leave the image and source together
Retain the obtained review images, official production case images, and original webpage video covers according to the source. Official interaction charts record the default configuration opened this time; Different inference files, tools, cost axes, and test versions must not be mixed and compared. GIFs retain the original files; video covers do not represent the full viewing or retesting of this article. Advertisements, author avatars, brand logos, and related article recommendation images are not included in this review catalog.
Official release page: 30 charts
Security tasks include research settings without production protection. Research environment scores do not equal online product access, nor do they equal daily success rates. Screenshots include complete charts; explanations below and test footnotes should be based on the original webpage.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Official release page: Interaction behavior comparison
This is an officially selected interaction example showing requirements clarification and collaboration differences, which cannot be used to calculate the general win rate.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Independent Institution First Release Test: Legacy Index
This is the old index chart from September 3, preserving the conclusions and criteria from that time; For the current ranking, please refer to the 4.2 map that follows, without cross-version comparison.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Independent Institutions New Version Test: Index 4.2

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Interactive reasoning: Operation mode and action efficiency

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Early Users: Game Creation Effects

Image source: Original author’s page; View original image or chart.
Official game development: concept, gameplay, and engineering validation

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
Official 3D production: models, materials, and engine effects

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.

Image source: Original author’s page; View original image or chart.
The four gadgets I’m working on
HyphenBox
Free large model interface radar, plus local unified routing.
- Current Status: Initial Build · Preview; For Apple chip Macs, self-signed, manual release is required for the first time
- Download and Update
LocalBrain
Turn your Mac into your private AI toolbox, integrating transcription, voiceover, raw images, video, and MCP.
- Current Status: Continuous iteration
- Download and Update
ScreenLex ScreenLex
Turn new words from local movie subtitles into a reviewable English vocabulary, fully offline.
- Current Status: Continuous iteration
- Download and Update
🧰 Tools I build
I maintain all of these tools myself. Preview builds are clearly labeled; the release pages are the source of truth for downloads, updates and known limits.
Info: HyphenBox Status: Official releases
A radar for free LLM APIs: availability is re-tested continuously, one local interface for all of them, and keys stay on your machine
Info: LocalBrain Status: Official releases
A multimodal MCP toolbox for local models: TTS, Whisper and video generation in one place
Info: ScreenLex Status: Official releases
Learn new words while you watch shows. Free, for Mac and Windows
Info: HyphenScreen Status: Official releases
Screen recording and smart editing in one: a DaVinci-style timeline, automatic redaction and a check of the finished video before export. Free
Quote: HyphenTech Make AI your superpower Local deployment · Free resources · Self-made software https://hyphentech.top
Late nights and burned API credits went in,a cup of tea comes back out — only if you feel like it.
Scan with WeChatPress and hold to save the image, then open it from your album in WeChat Scan


Comments
Loading comments…