AI 产品经理 · 能独立交付的 AI BuilderAI Product Manager · Builder who ships end to end

张潆木Yingmu ZhangYingmu Zhang张潆木

2 年荷兰 B2B SaaS 研发,1.5 年一人公司Two years building B2B SaaS in Rotterdam. Eighteen months running a one-person company.

一个 iOS 产品上架、一个线上工具、多次 AI 产品可行性验证One iOS app shipped, one web tool live, several AI feasibility probes run to a verdict.

wawamu2201@gmail.com·微信 wawawazymWeChat wawawazym

229iOS 应用真实测试用户
2,494 次会话 · 0 崩溃
real TestFlight testers
2,494 sessions · 0 crashes
2.5万25K自有内容账号粉丝
产品冷启动渠道
followers on my own channel
cold-start channel for my products
128篇文献全量结构化抽取
检索 6,471 篇
papers fully extracted into fields
6,471 retrieved
87%前端测试覆盖率
由 60% 推动至此
frontend test coverage
driven up from 60%
Capability

能力

Capability

AI 产品

AI Product

Vibecoding · RAG 方案论证与边界判断 · Agent 编排取舍 · 检索与抽取管线设计 · 标注 schema 与一致率评测 · badcase 归因 · 生成措辞与安全边界 · human-in-the-loop 审核 · Prompt 工程 · LangChain / LangGraph

Vibecoding · RAG solution reasoning and boundary judgement · agent orchestration trade-offs · retrieval and extraction pipeline design · annotation schema and agreement evaluation · badcase attribution · generation wording and safety boundaries · human-in-the-loop review · prompt engineering · LangChain / LangGraph

产品交付

Product Delivery

PRD / TDD 撰写 · MVP 范围裁剪与取舍论证 · 埋点与指标口径设计 · 内测组织与反馈闭环 · App Store 发布

PRD / TDD writing · MVP scoping and trade-off argumentation · event tracking and metric definitions · beta programmes and feedback loops · App Store release

技术实现

Engineering

TypeScript · Vue 3 / React / Next.js · React Native · Node.js · Python · Docker / AWS CI-CD · 自动化测试 · Claude Code Skill / Hook / 自定义 Command

TypeScript · Vue 3 / React / Next.js · React Native · Node.js · Python · Docker / AWS CI-CD · automated testing · Claude Code Skills / Hooks / custom commands

协作与语言

Collaboration & Languages

Agile / Scrum 跨职能协作 · 中文母语 · 英语可作为工作语言 · 英国本科 + 荷兰硕士 + 2 年英语工作环境

Agile / Scrum cross-functional collaboration · native Chinese · English as a working language · UK bachelor’s + Dutch master’s + two years in an English-speaking workplace

Work

产品与项目

Products & Projects

WorryBucket · 定点焦虑桶WorryBucket

iOS · 已提交审核 · 10/15 发布iOS · in review · ships Oct 15

把焦虑的「记录」与「处理」拆成两段,各自服务不同的认知状态:收集时三秒一行字、不问感受不打分;到自己设定的时间再一次处理整桶。一人完成产品定义、交互设计、开发与提审。

Separates recording anxiety from processing it, because the two serve different cognitive states: three seconds and one line when you capture — no ratings, no questions about how you feel; then the whole bucket is worked through once, at a time you set. Product definition, interaction design, build and App Store submission, all solo.

229TestFlight 测试者TestFlight testers
2,494会话sessions
0崩溃crashes
¥0获客成本acquisition cost
Three deliberate refusals: no streaks, no unread badges, no comparison between users. Each of them would lift DAU; each of them would harm these users. They are written into the design system as prohibitions, not as “not yet”.
Nothing leaves the device: no account, no analytics, no network calls. App Privacy label reads Data Not Collected.
Constraints encoded for the agent, not recited at it: the hard rules in the PRD / TDD are written as three Claude Code Skills (AI guidance layer / capture interaction / state machine), so they take effect automatically during implementation. The first red line in the AI layer — no keys on the client, every call through a backend proxy — is an architectural spec inside the Skill, not a verbal agreement. Four product red lines are separately enforced as a pre-commit hook: a hit blocks the commit. Sixteen weeks, zero violations.
三个主动拒绝:不做连胜、不做未读红点、不做用户间比较——每一个都会降低日活,但对这批用户都是伤害。写进设计系统当禁止项,不是「暂时没做」。
零数据出网:无账号、无分析、无网络调用,App Privacy 标记 Data Not Collected。
把约束编码给 Agent,而不是叮嘱它:PRD / TDD 里的硬规则写成 3 个 Claude Code Skill(AI 引导层 / 收集交互 / 状态机),实现阶段自动生效。比如 AI 层的第一条红线——密钥不落客户端、调用全部走后端代理——是 Skill 里的架构规格,不是口头约定。四条产品红线另做成提交前钩子,命中即阻断,16 周零违规。

循证类 AI 产品可行性验证

Evidence-based AI product · feasibility probes

两次 · 按判据终止two runs · terminated on pre-set criteria

从非结构化英文文献抽结构化字段,并验证抽得对不对——与合同审查、理赔材料解析、工单结构化是同一类问题。数据源 PubMed / PMC,自建抽取管线与人工标注工作台。

Pulling structured fields out of unstructured English literature — and verifying whether the extraction is actually correct. The same class of problem as contract review, claims-document parsing and ticket structuring. Source: PubMed / PMC, with a self-built extraction pipeline and human annotation workbench.

6,471篇检索papers retrieved
128篇全量抽取fully extracted
84.9%条目级一致率item-level agreement
Model evaluation — a live case of proxy-metric decoupling. Schema compliance came in at 95.3%; substantive accuracy was 40%. The main decision rule turned out to be unexecutable on two-thirds of the samples because the preconditions were simply absent from the source — and every one of those cases passed automated validation with zero alerts (silent failure). The transferable conclusion: an eval set has to measure task completion, not output conformance; format validators are structurally blind to this class of error.
Annotation governance. Field-level schema and decision thresholds were pre-registered — fixed before the run and never moved afterwards, so the criteria could not drift toward the result. Every human judgement was made by me: 195 relevance reads plus 20 targeted verifications, with five-category badcase attribution. Independent review was made structural rather than procedural: the adjudicator cannot see the model's output at the code level, which removes anchoring by mechanism instead of by discipline. Agreement is reported per field, never as a single blended average — an aggregate number hides exactly the fields that are failing.
Boundary conditions and safety. Generation wording was constrained by a field-combination → permitted-phrasing lookup table, run as a deterministic pre-check before generation rather than as a filter after it; any phrasing that predicts an individual outcome is unconditionally blocked. Body copy is model-generated and passes human-in-the-loop review before publication.
Technology selection and capability boundaries. RAG and multi-agent orchestration were both evaluated, both rejected, and both decisions documented. Request-time RAG would let every output bypass human review, which is not acceptable in a high-risk health-information domain; the pipeline runs at build time instead, with full human review before anything ships. A vector store over abstracts would amplify “a significant association was found” and suppress “confounders were not ruled out” — reproducing precisely the misreading this product exists to correct, namely taking a correlation from an observational study for causation. And standard retrieval metrics cannot detect it: they measure whether what was retrieved is relevant, not whether what was missed would overturn the answer.
模型评测 —— 一次代理指标失效的实测。schema 合规率 95.3%,实质正确率 40%。主判定规则在三分之二样本上无法执行,原因不是模型能力不足,而是判定所需的前提数据在原文里根本不存在;而这批样本全量通过自动校验、零告警(静默失效)。可迁移的结论:评测集要测任务完成度,不能只测输出规范性——格式校验层对这类错误是结构性失明的。
标注治理。字段级 schema 与判定阈值事前预注册:写死后再跑,不回头改线,避免判据向结果漂移。全部人工判定独立完成——195 篇相关性判读 + 20 篇定向核对,配 badcase 五类归因。独立复核做成结构性的而非流程性的:判定者在代码层面看不到模型输出,用机制而不是自律排除锚定。一致率分字段报出,不报总均值——总均值恰好会盖住那几个正在出问题的字段。
边界条件与安全。生成措辞由「字段组合 → 允许措辞」映射表约束,作为生成前的确定性预检,而不是生成后的过滤;任何指向个体未来的预测句式无条件禁止。正文由模型生成,经 human-in-the-loop 审核后才进入发布。
选型与能力边界。RAG 与多 Agent 编排都评估过、都否决了,决策记录留档。请求期 RAG 会让每条输出绕过人工复核,在高风险健康议题上不可接受,改为管线在构建期跑、产物全量人审后上线。摘要层向量库会放大「有显著相关」、隐去「未排除混杂」,恰好复制这个产品要纠正的核心误读——把观察性研究的相关性当成因果。而常规检索指标测不出这一点:它们测「召回的相不相关」,不测「没召回的会不会推翻答案」。

为需要日更图文的创作者解决「每天从空白画布重新排版」。浏览器内打开即用、无需注册,数据只存本机。

Solves “starting from a blank canvas every single day” for creators who publish daily. Opens straight in the browser, no sign-up, data stays on the device.

19周weeks
93张工单tickets
−77%首屏阻塞资源render-blocking assets
Where AI belongs here was argued for, not bolted on. The real demand runs image → words, not words → image: you start from a picture you took or saved, and the language comes out of it. That needs multimodal visual understanding — a rule engine cannot do it — so here AI is genuinely the only option. Which let me draw the boundary once and for all: AI decides where content comes from, composition rules decide how form is generated — the nine principles of graphic composition are all deterministic geometry (loops, interpolation, polar transforms, density fields) and never touch a model. Every later feature already knows which side it belongs on. The same line explains why the first two versions carried no AI: not because it was hard, but because nothing there genuinely required it.
Choosing the model overturned my own selection criterion. Benchmarking four Chinese vision models for image-to-copy, I had planned to filter on cost; in practice cost turned out to be a dead criterion — ¥0.003 per call separated the most from the least expensive, which eliminates nobody. I switched the criterion to whether structured output has an architectural guarantee. On architecture: an OpenAI-compatible protocol plus three environment variables isolate the vendor, so nothing is welded to one provider. Within a month of that research all three vendors changed models and prices — which validated the decision from the other direction.
Harness engineering: constrain with mechanism, don't pray with prompts. “Please run lint before committing” in CLAUDE.md is a suggestion the model is free to ignore. Rewritten as three hooks (block commits on the trunk / auto-format on write / force a typecheck at stop) plus two custom commands, the delivery process went from “remember to” to “you cannot get past it”.
“Open it and use it” is the promise this product makes, and first paint is where that promise is kept or broken. A tool with no login wall is judged entirely on what the first load shows — and 1.1MB of render-blocking assets means the promise does not hold for anyone on a weak connection. The valuable part was not the fix but declining the obvious one: I was about to build a more elaborate lazy-loading scheme, and diagnosing first showed the root cause was a font preload call passing no character set, which defeated range-based on-demand loading entirely. One call changed: 1.1MB → 255KB. Measuring first saved an entire piece of engineering that would have been wasted.
AI 用在哪,是论证出来的,不是加上去的。真实需求的流向是「图 → 词」而不是「词 → 图」——人先有一张图(拍的、存的、看到的),语言内容从图里发散出来;这需要多模态视觉理解,规则引擎做不了,所以这里的 AI 是非它不可。由此把边界一次画清:AI 管内容从哪来,构成法则管形式怎么生成——九种平面构成法则全是确定性几何算法(循环 / 插值 / 极坐标变换 / 密度场),不走模型。后面每个新功能该归哪边,都有现成答案。同一条线也解释了前两个版本为什么没上 AI:不是做不出来,是当时没有非它不可的场景。
AI 能力选型推翻了我自己原定的筛选维度。为「看图出词」横评四个国产视觉模型,原计划按成本筛,实测下来成本这条筛选条件是失效的——最贵与最便宜相差 ¥0.003/次,一个候选都筛不掉。改以「结构化输出有没有架构级保障」定选。架构上用 OpenAI 兼容协议 + 三个环境变量隔离供应商,不焊死任何一家。
Harness Engineering:用机制约束,不用提示词祈祷。「提交前请跑 lint」写进 CLAUDE.md 只是建议,模型可以忽略;改写成 3 个 hook(主干分支拦截提交 / 写入后自动格式化 / 收尾强制类型检查)与 2 个自定义 command 之后,交付流程从「记得做」变成「做不到就过不去」。
「打开即用」是这个产品的卖点,而它的兑现点就在首屏。工具类产品没有登录墙拦着,用户第一次打开看到什么就是全部印象——阻塞渲染资源 1.1MB,意味着这个承诺在弱网用户那里根本不成立。但真正值钱的不是修好了,是没有按第一反应去修:我本来要上更复杂的懒加载方案,先做定位才发现根因只是字体预加载调用未传字符集、把按需分片加载全量废掉了。改这一处,1.1MB → 255KB。先量再动手,省掉的是一整套本来会白做的工程。
Experience

经历

Experience

自媒体创业与独立产品开发

Content Business & Independent Product Development

一人公司 · 产品开发 / 内容运营 / 付费咨询
One-person company · product · content · paid consulting
  • 完整走通从产品定义、开发交付到上架发布的全流程,用户全部来自自有渠道
  • Ran the full path from product definition through build to store release; every user came from a channel I own
  • 运营 2.5 万粉知识账号:选题即需求验证,同时是自有产品的冷启动渠道
  • Run a 25K-follower knowledge account: topic selection doubles as demand validation and as the cold-start channel for my own products

前端开发工程师

Frontend Engineer

4Shipping · 鹿特丹,荷兰 · 内河航运 B2B SaaS
4Shipping · Rotterdam, Netherlands · inland-shipping B2B SaaS
  • 多角色产品中处理「同一功能对不同角色含义不同」带来的设计约束
  • Worked the design constraints of a multi-role product, where one feature means different things to shipowners, cargo owners and carriers
  • 主导复杂动态表格的设计与迭代,用虚拟滚动与分批加载重构渲染方案
  • Led the design and iteration of complex dynamic tables; rebuilt rendering with virtual scrolling and batched loading
  • 推动测试自动化与 TDD,前端测试覆盖率由 60% 提升至 87%
  • Drove test automation and TDD; frontend test coverage from 60% to 87%
  • 参与前端 v2 → v3 重大升级迁移
  • Took part in the major v2 → v3 frontend migration

全栈项目助教

Full-stack Teaching Assistant

Matrix Master · 鹿特丹
Matrix Master · Rotterdam
  • 带零基础与转行学员完成全栈项目实践,核心是把技术概念翻译成不同基础的人能听懂的语言
  • Guided beginners and career-changers through full-stack projects; the real work was translating technical concepts into language each person could actually follow
Education

教育

Education

新媒体与数字文化 · 硕士

New Media & Digital Culture · MA

阿姆斯特丹大学 · 选修 Python、数据分析、移动应用软件研究
University of Amsterdam · electives: Python, data analysis, app studies

国际传播学 · 学士

International Communications · BA

诺丁汉大学 · 选修社交媒体研究、Web 基础
University of Nottingham · electives: social media research, web fundamentals
Contact

欢迎聊聊

Let's talk