iOS · 已提交审核 · 10/15 发布iOS · in review · ships Oct 15
把焦虑的「记录」与「处理」拆成两段,各自服务不同的认知状态:收集时三秒一行字、不问感受不打分;到自己设定的时间再一次处理整桶。一人完成产品定义、交互设计、开发与提审。
Separates recording anxiety from processing it, because the two serve different cognitive states: three seconds and one line when you capture — no ratings, no questions about how you feel; then the whole bucket is worked through once, at a time you set. Product definition, interaction design, build and App Store submission, all solo.
229TestFlight 测试者TestFlight testers
2,494会话sessions
0崩溃crashes
¥0获客成本acquisition cost
Three deliberate refusals: no streaks, no unread badges, no comparison between users. Each of them would lift DAU; each of them would harm these users. They are written into the design system as prohibitions, not as “not yet”.
Nothing leaves the device: no account, no analytics, no network calls. App Privacy label reads Data Not Collected.
Constraints encoded for the agent, not recited at it: the hard rules in the PRD / TDD are written as three Claude Code Skills (AI guidance layer / capture interaction / state machine), so they take effect automatically during implementation. The first red line in the AI layer — no keys on the client, every call through a backend proxy — is an architectural spec inside the Skill, not a verbal agreement. Four product red lines are separately enforced as a pre-commit hook: a hit blocks the commit. Sixteen weeks, zero violations.
三个主动拒绝:不做连胜、不做未读红点、不做用户间比较——每一个都会降低日活,但对这批用户都是伤害。写进设计系统当禁止项,不是「暂时没做」。
零数据出网:无账号、无分析、无网络调用,App Privacy 标记 Data Not Collected。
把约束编码给 Agent,而不是叮嘱它:PRD / TDD 里的硬规则写成 3 个 Claude Code Skill(AI 引导层 / 收集交互 / 状态机),实现阶段自动生效。比如 AI 层的第一条红线——密钥不落客户端、调用全部走后端代理——是 Skill 里的架构规格,不是口头约定。四条产品红线另做成提交前钩子,命中即阻断,16 周零违规。
循证类 AI 产品可行性验证
Evidence-based AI product · feasibility probes
两次 · 按判据终止two runs · terminated on pre-set criteria
从非结构化英文文献抽结构化字段,并验证抽得对不对——与合同审查、理赔材料解析、工单结构化是同一类问题。数据源 PubMed / PMC,自建抽取管线与人工标注工作台。
Pulling structured fields out of unstructured English literature — and verifying whether the extraction is actually correct. The same class of problem as contract review, claims-document parsing and ticket structuring. Source: PubMed / PMC, with a self-built extraction pipeline and human annotation workbench.
6,471篇检索papers retrieved
128篇全量抽取fully extracted
84.9%条目级一致率item-level agreement
The single most valuable finding: format-compliance came in at 95.3%, but the main decision rule was physically unexecutable on two-thirds of the samples — substantive accuracy fell to 40%, and the format-validation layer caught none of it.
The criteria for the two runs were designed in sequence: the second replaced exact thresholds with coarse bands, reported agreement per field instead of in aggregate, and made independent review structural — the adjudicator cannot see the model's output at the code level.
RAG and multi-agent orchestration were both evaluated, both rejected, and both reasons written down. Request-time RAG was rejected on safety: this is a high-risk health-information domain, and generating at request time means every output bypasses human review. The pipeline runs at build time instead, and nothing is published until a human has reviewed all of it. A vector store built over abstracts was rejected on epistemics: it will faithfully retrieve “significantly associated” and never retrieve “co-occurring factors were not ruled out”, because the latter is not in the index at all. The effect is systematic: it amplifies “a significant association was found” and suppresses “confounders were not ruled out” — which is precisely the misreading evidence-based health writing exists to correct: taking a correlation from an observational study for causation, when the confounders were never excluded in the first place. Build the retrieval layer that way and the product manufactures the very thing it was built to counter. And standard retrieval metrics cannot see this: they measure whether what was retrieved is relevant, not whether what was missed would overturn the answer.
The annotation system is my own: field-level schema and decision thresholds pre-registered — fixed before the run and never moved afterwards — with every human judgement made by me: 195 relevance reads plus 20 targeted verifications.
最值钱的一条发现:模型格式合规率 95.3%,但主判定规则在三分之二的样本上物理上无法执行——实质正确率掉到 40%,而格式校验层一个都查不出来。
两次的判据设计前后相继:第二次把精确阈值换成粗档、一致率改为分字段报、独立复核做成结构性的(判定者在代码层面看不到模型输出)。
RAG 与多 Agent 编排都评估过,也都否决了,理由都写下来了。请求期 RAG 因安全否决:这是高风险健康议题,运行时生成意味着每一条输出都绕过人工复核。改为管线在构建期跑、产物全量人工审核后才上线,把不确定性挡在发布之前。在摘要上建向量库因认识论否决——它会忠实召回「显著相关」,却永远召不回「共现因素未被排除」,因为后者根本不在索引里。结果是它会系统性地放大「有显著相关」、隐去「未排除混杂」,而这恰好就是循证健康科普要纠正的核心误读:把观察性研究里的相关性当成因果,而混杂因素其实从未被排除。检索层如果这样建,产品就会亲手复制它本来要对抗的东西。而常规检索指标测不出这一点:它们测「召回的相不相关」,不测「没召回的会不会推翻答案」。
标注体系是自己搭的:字段级 schema 与判定阈值事前预注册(写死后再跑,不回头改线),全部人工判定由我完成——195 篇相关性判读 + 20 篇定向核对。
为需要日更图文的创作者解决「每天从空白画布重新排版」。浏览器内打开即用、无需注册,数据只存本机。
Solves “starting from a blank canvas every single day” for creators who publish daily. Opens straight in the browser, no sign-up, data stays on the device.
19周weeks
93张工单tickets
−77%首屏阻塞资源render-blocking assets
Where AI belongs here was argued for, not bolted on. The real demand runs image → words, not words → image: you start from a picture you took or saved, and the language comes out of it. That needs multimodal visual understanding — a rule engine cannot do it — so here AI is genuinely the only option. Which let me draw the boundary once and for all: AI decides where content comes from, composition rules decide how form is generated — the nine principles of graphic composition are all deterministic geometry (loops, interpolation, polar transforms, density fields) and never touch a model. Every later feature already knows which side it belongs on. The same line explains why the first two versions carried no AI: not because it was hard, but because nothing there genuinely required it.
Choosing the model overturned my own selection criterion. Benchmarking four Chinese vision models for image-to-copy, I had planned to filter on cost; in practice cost turned out to be a dead criterion — ¥0.003 per call separated the most from the least expensive, which eliminates nobody. I switched the criterion to whether structured output has an architectural guarantee. On architecture: an OpenAI-compatible protocol plus three environment variables isolate the vendor, so nothing is welded to one provider. Within a month of that research all three vendors changed models and prices — which validated the decision from the other direction.
Harness engineering: constrain with mechanism, don't pray with prompts. “Please run lint before committing” in CLAUDE.md is a suggestion the model is free to ignore. Rewritten as three hooks (block commits on the trunk / auto-format on write / force a typecheck at stop) plus two custom commands, the delivery process went from “remember to” to “you cannot get past it”.
“Open it and use it” is the promise this product makes, and first paint is where that promise is kept or broken. A tool with no login wall is judged entirely on what the first load shows — and 1.1MB of render-blocking assets means the promise does not hold for anyone on a weak connection. The valuable part was not the fix but declining the obvious one: I was about to build a more elaborate lazy-loading scheme, and diagnosing first showed the root cause was a font preload call passing no character set, which defeated range-based on-demand loading entirely. One call changed: 1.1MB → 255KB. Measuring first saved an entire piece of engineering that would have been wasted.
AI 用在哪,是论证出来的,不是加上去的。真实需求的流向是「图 → 词」而不是「词 → 图」——人先有一张图(拍的、存的、看到的),语言内容从图里发散出来;这需要多模态视觉理解,规则引擎做不了,所以这里的 AI 是非它不可。由此把边界一次画清:AI 管内容从哪来,构成法则管形式怎么生成——九种平面构成法则全是确定性几何算法(循环 / 插值 / 极坐标变换 / 密度场),不走模型。后面每个新功能该归哪边,都有现成答案。同一条线也解释了前两个版本为什么没上 AI:不是做不出来,是当时没有非它不可的场景。
AI 能力选型推翻了我自己原定的筛选维度。为「看图出词」横评四个国产视觉模型,原计划按成本筛,实测下来成本这条筛选条件是失效的——最贵与最便宜相差 ¥0.003/次,一个候选都筛不掉。改以「结构化输出有没有架构级保障」定选。架构上用 OpenAI 兼容协议 + 三个环境变量隔离供应商,不焊死任何一家。
Harness Engineering:用机制约束,不用提示词祈祷。「提交前请跑 lint」写进 CLAUDE.md 只是建议,模型可以忽略;改写成 3 个 hook(主干分支拦截提交 / 写入后自动格式化 / 收尾强制类型检查)与 2 个自定义 command 之后,交付流程从「记得做」变成「做不到就过不去」。
「打开即用」是这个产品的卖点,而它的兑现点就在首屏。工具类产品没有登录墙拦着,用户第一次打开看到什么就是全部印象——阻塞渲染资源 1.1MB,意味着这个承诺在弱网用户那里根本不成立。但真正值钱的不是修好了,是没有按第一反应去修:我本来要上更复杂的懒加载方案,先做定位才发现根因只是字体预加载调用未传字符集、把按需分片加载全量废掉了。改这一处,1.1MB → 255KB。先量再动手,省掉的是一整套本来会白做的工程。