iOS · 已提交审核 · 10/15 发布iOS · in review · ships Oct 15
把焦虑的「记录」与「处理」拆成两段,各自服务不同的认知状态:收集时三秒一行字、不问感受不打分;到自己设定的时间再一次处理整桶。一人完成产品定义、交互设计、开发与提审。
Separates recording anxiety from processing it, because the two serve different cognitive states: three seconds and one line when you capture — no ratings, no questions about how you feel; then the whole bucket is worked through once, at a time you set. Product definition, interaction design, build and App Store submission, all solo.
229TestFlight 测试者TestFlight testers
2,494会话sessions
0崩溃crashes
¥0获客成本acquisition cost
Three deliberate refusals: no streaks, no unread badges, no comparison between users. Each of them would lift DAU; each of them would harm these users. They are written into the design system as prohibitions, not as “not yet”.
Nothing leaves the device: no account, no analytics, no network calls. App Privacy label reads Data Not Collected.
Constraints encoded for the agent, not recited at it: the hard rules in the PRD / TDD are written as three Claude Code Skills (AI guidance layer / capture interaction / state machine), so they take effect automatically during implementation. The first red line in the AI layer — no keys on the client, every call through a backend proxy — is an architectural spec inside the Skill, not a verbal agreement. Four product red lines are separately enforced as a pre-commit hook: a hit blocks the commit. Sixteen weeks, zero violations.
三个主动拒绝:不做连胜、不做未读红点、不做用户间比较——每一个都会降低日活,但对这批用户都是伤害。写进设计系统当禁止项,不是「暂时没做」。
零数据出网:无账号、无分析、无网络调用,App Privacy 标记 Data Not Collected。
把约束编码给 Agent,而不是叮嘱它:PRD / TDD 里的硬规则写成 3 个 Claude Code Skill(AI 引导层 / 收集交互 / 状态机),实现阶段自动生效。比如 AI 层的第一条红线——密钥不落客户端、调用全部走后端代理——是 Skill 里的架构规格,不是口头约定。四条产品红线另做成提交前钩子,命中即阻断,16 周零违规。
循证类 AI 产品可行性验证
Evidence-based AI product · feasibility probes
两次 · 按判据终止two runs · terminated on pre-set criteria
从非结构化英文文献抽结构化字段,并验证抽得对不对——与合同审查、理赔材料解析、工单结构化是同一类问题。数据源 PubMed / PMC,自建抽取管线与人工标注工作台。
Pulling structured fields out of unstructured English literature — and verifying whether the extraction is actually correct. The same class of problem as contract review, claims-document parsing and ticket structuring. Source: PubMed / PMC, with a self-built extraction pipeline and human annotation workbench.
6,471篇检索papers retrieved
128篇全量抽取fully extracted
84.9%条目级一致率item-level agreement
Model evaluation — a live case of proxy-metric decoupling. Schema compliance came in at 95.3%; substantive accuracy was 40%. The main decision rule turned out to be unexecutable on two-thirds of the samples because the preconditions were simply absent from the source — and every one of those cases passed automated validation with zero alerts (silent failure). The transferable conclusion: an eval set has to measure task completion, not output conformance; format validators are structurally blind to this class of error.
Annotation governance. Field-level schema and decision thresholds were pre-registered — fixed before the run and never moved afterwards, so the criteria could not drift toward the result. Every human judgement was made by me: 195 relevance reads plus 20 targeted verifications, with five-category badcase attribution. Independent review was made structural rather than procedural: the adjudicator cannot see the model's output at the code level, which removes anchoring by mechanism instead of by discipline. Agreement is reported per field, never as a single blended average — an aggregate number hides exactly the fields that are failing.
Boundary conditions and safety. Generation wording was constrained by a field-combination → permitted-phrasing lookup table, run as a deterministic pre-check before generation rather than as a filter after it; any phrasing that predicts an individual outcome is unconditionally blocked. Body copy is model-generated and passes human-in-the-loop review before publication.
Technology selection and capability boundaries. RAG and multi-agent orchestration were both evaluated, both rejected, and both decisions documented. Request-time RAG would let every output bypass human review, which is not acceptable in a high-risk health-information domain; the pipeline runs at build time instead, with full human review before anything ships. A vector store over abstracts would amplify “a significant association was found” and suppress “confounders were not ruled out” — reproducing precisely the misreading this product exists to correct, namely taking a correlation from an observational study for causation. And standard retrieval metrics cannot detect it: they measure whether what was retrieved is relevant, not whether what was missed would overturn the answer.
模型评测 —— 一次代理指标失效的实测。schema 合规率 95.3%,实质正确率 40%。主判定规则在三分之二样本上无法执行,原因不是模型能力不足,而是判定所需的前提数据在原文里根本不存在;而这批样本全量通过自动校验、零告警(静默失效)。可迁移的结论:评测集要测任务完成度,不能只测输出规范性——格式校验层对这类错误是结构性失明的。
标注治理。字段级 schema 与判定阈值事前预注册:写死后再跑,不回头改线,避免判据向结果漂移。全部人工判定独立完成——195 篇相关性判读 + 20 篇定向核对,配 badcase 五类归因。独立复核做成结构性的而非流程性的:判定者在代码层面看不到模型输出,用机制而不是自律排除锚定。一致率分字段报出,不报总均值——总均值恰好会盖住那几个正在出问题的字段。
边界条件与安全。生成措辞由「字段组合 → 允许措辞」映射表约束,作为生成前的确定性预检,而不是生成后的过滤;任何指向个体未来的预测句式无条件禁止。正文由模型生成,经 human-in-the-loop 审核后才进入发布。
选型与能力边界。RAG 与多 Agent 编排都评估过、都否决了,决策记录留档。请求期 RAG 会让每条输出绕过人工复核,在高风险健康议题上不可接受,改为管线在构建期跑、产物全量人审后上线。摘要层向量库会放大「有显著相关」、隐去「未排除混杂」,恰好复制这个产品要纠正的核心误读——把观察性研究的相关性当成因果。而常规检索指标测不出这一点:它们测「召回的相不相关」,不测「没召回的会不会推翻答案」。
为需要日更图文的创作者解决「每天从空白画布重新排版」。浏览器内打开即用、无需注册,数据只存本机。
Solves “starting from a blank canvas every single day” for creators who publish daily. Opens straight in the browser, no sign-up, data stays on the device.
19周weeks
93张工单tickets
−77%首屏阻塞资源render-blocking assets
Where AI belongs here was argued for, not bolted on. The real demand runs image → words, not words → image: you start from a picture you took or saved, and the language comes out of it. That needs multimodal visual understanding — a rule engine cannot do it — so here AI is genuinely the only option. Which let me draw the boundary once and for all: AI decides where content comes from, composition rules decide how form is generated — the nine principles of graphic composition are all deterministic geometry (loops, interpolation, polar transforms, density fields) and never touch a model. Every later feature already knows which side it belongs on. The same line explains why the first two versions carried no AI: not because it was hard, but because nothing there genuinely required it.
Choosing the model overturned my own selection criterion. Benchmarking four Chinese vision models for image-to-copy, I had planned to filter on cost; in practice cost turned out to be a dead criterion — ¥0.003 per call separated the most from the least expensive, which eliminates nobody. I switched the criterion to whether structured output has an architectural guarantee. On architecture: an OpenAI-compatible protocol plus three environment variables isolate the vendor, so nothing is welded to one provider. Within a month of that research all three vendors changed models and prices — which validated the decision from the other direction.
Harness engineering: constrain with mechanism, don't pray with prompts. “Please run lint before committing” in CLAUDE.md is a suggestion the model is free to ignore. Rewritten as three hooks (block commits on the trunk / auto-format on write / force a typecheck at stop) plus two custom commands, the delivery process went from “remember to” to “you cannot get past it”.
“Open it and use it” is the promise this product makes, and first paint is where that promise is kept or broken. A tool with no login wall is judged entirely on what the first load shows — and 1.1MB of render-blocking assets means the promise does not hold for anyone on a weak connection. The valuable part was not the fix but declining the obvious one: I was about to build a more elaborate lazy-loading scheme, and diagnosing first showed the root cause was a font preload call passing no character set, which defeated range-based on-demand loading entirely. One call changed: 1.1MB → 255KB. Measuring first saved an entire piece of engineering that would have been wasted.
AI 用在哪,是论证出来的,不是加上去的。真实需求的流向是「图 → 词」而不是「词 → 图」——人先有一张图(拍的、存的、看到的),语言内容从图里发散出来;这需要多模态视觉理解,规则引擎做不了,所以这里的 AI 是非它不可。由此把边界一次画清:AI 管内容从哪来,构成法则管形式怎么生成——九种平面构成法则全是确定性几何算法(循环 / 插值 / 极坐标变换 / 密度场),不走模型。后面每个新功能该归哪边,都有现成答案。同一条线也解释了前两个版本为什么没上 AI:不是做不出来,是当时没有非它不可的场景。
AI 能力选型推翻了我自己原定的筛选维度。为「看图出词」横评四个国产视觉模型,原计划按成本筛,实测下来成本这条筛选条件是失效的——最贵与最便宜相差 ¥0.003/次,一个候选都筛不掉。改以「结构化输出有没有架构级保障」定选。架构上用 OpenAI 兼容协议 + 三个环境变量隔离供应商,不焊死任何一家。
Harness Engineering:用机制约束,不用提示词祈祷。「提交前请跑 lint」写进 CLAUDE.md 只是建议,模型可以忽略;改写成 3 个 hook(主干分支拦截提交 / 写入后自动格式化 / 收尾强制类型检查)与 2 个自定义 command 之后,交付流程从「记得做」变成「做不到就过不去」。
「打开即用」是这个产品的卖点,而它的兑现点就在首屏。工具类产品没有登录墙拦着,用户第一次打开看到什么就是全部印象——阻塞渲染资源 1.1MB,意味着这个承诺在弱网用户那里根本不成立。但真正值钱的不是修好了,是没有按第一反应去修:我本来要上更复杂的懒加载方案,先做定位才发现根因只是字体预加载调用未传字符集、把按需分片加载全量废掉了。改这一处,1.1MB → 255KB。先量再动手,省掉的是一整套本来会白做的工程。