03 来源索引
返回 AI Harness
来源优先级
- OpenAI / Anthropic 官方工程博客、开发者文档、官方 guide。
- Google / Microsoft 等大厂官方云平台和工程博客。
- LangSmith / Langfuse / Braintrust 等 eval/observability 工具官方资料。
- 二手博客、awesome list、社区帖只做线索,不做主证据。
OpenAI
| 来源 | 主题 | 可用结论 |
|---|---|---|
| A practical guide to building agents | agent foundations、orchestration、guardrails | 可靠 agent 需要模型、工具、结构化指令、护栏和渐进式部署。 |
| New tools for building agents | Responses API、Agents SDK、tools、handoffs、guardrails、tracing | OpenAI 把 tools、handoffs、guardrails、tracing 作为 agent 应用 primitives。 |
| Safety in building agents | input guardrails、PII、jailbreak、trace graders、tool approvals | 对不可信输入用结构化字段、隔离、tool approval 和 human approval 降低风险。 |
| Evaluate agent workflows | traces、graders、datasets、eval runs | 从 trace 调试行为,再沉淀 dataset 和 eval runs 做可重复评估。 |
| How to implement LLM guardrails | guardrail cookbook | 可作为实现输入/输出 guardrails 的代码参考。 |
| Building Governed AI Agents | governed agents、scaffolding | 作为 governance 和 tool policy 的补充参考;优先级低于官方 guide。 |
Anthropic
| 来源 | 主题 | 可用结论 |
|---|---|---|
| Trustworthy agents in practice | model / harness / tools / environment | Anthropic 明确把 harness 定义为模型运行其上的 instructions 和 guardrails。 |
| Effective context engineering for AI agents | context、tools、examples、message history、runtime retrieval | context engineering 是管理完整 context state,不是只改 prompt。 |
| Writing effective tools for agents | tool design、high-signal output、eval-driven iteration | 工具要清晰、少重叠、返回高信号结果,并通过 eval/transcript 迭代。 |
| Building Effective AI Agents | workflow patterns、simplicity、transparency、ACI | 简单、透明、良好的 agent-computer interface 是可靠 agent 的基础。 |
| Demystifying evals for AI agents | agent harness、eval suite、evaluation harness、outcome | 评估 agent 时是在评估 model + harness;要看 trajectory 和 final outcome。 |
| Harness design for long-running application development | long-running agents、context reset、handoff、self-evaluation | 长任务需要 context reset、structured handoff 和独立 evaluator。 |
| How we built our multi-agent research system | multi-agent、tool design、context/memory | 可作为复杂 agent 的旁证,尤其是工具描述、context 管理和子代理 handoff。 |
| 来源 | 主题 | 可用结论 |
|---|---|---|
| A dev’s guide to production-ready AI agents | production agents、trajectory eval、sandbox/canary/prod | agent 评估要看完整决策路径,部署要分 sandbox、canary、production。 |
| Startup technical guide: AI agents | ADK、observability、evaluation、context management | ADK 把 context management、evaluation、observability、containerization、multi-agent composition 放进生产能力。 |
| Agent Factory Recap: Securing AI Agents in Production | network isolation、logging、tool safeguards | 安全 harness 需要网络隔离、失败日志、工具保护和审计线索。 |
Microsoft
| 来源 | 主题 | 可用结论 |
|---|---|---|
| Agent Factory: Top 5 agent observability best practices for reliable AI | Azure AI Foundry、evals、red teaming、CI/CD、dashboards | observability 要覆盖开发、CI/CD 和生产,并纳入 safety/governance。 |
| Agent evaluators for generative AI | agent evaluators、workflow quality | 生产级 agent 需要评估每一步 workflow,不只是 final output。 |
| Observability for AI Systems | multi-turn trace、evaluation、governance | AI observability 的关联范围应覆盖完整 conversation / memory lifecycle。 |
Agent Search and Web Retrieval
| Source | Topic | Usable conclusion |
|---|---|---|
| Parallel Search · pricing | independent index, excerpts, Search, Task, monitoring | Parallel is designed for agent-native research; Search and high-depth Task processors are separate cost and latency classes. |
| Exa Search API · pricing | semantic search, specialized indexes, structured output | Exa is strongest as a semantic discovery layer for entities, research, and code rather than a traditional consumer SERP. |
| Tavily API · credits | search, extract, crawl, map, research | Tavily offers the broadest integrated retrieval workflow, but endpoint depth and credit consumption must be included in cost evaluation. |
| Brave Search API · pricing | independent index, web verticals, LLM Context | Brave is a strong independent raw-search provider and fallback when avoiding correlated Google/Bing dependencies matters. |
| You.com Search · billing | search, contents, managed research | Search, page contents, and multi-tier research are separate products and should be benchmarked and budgeted independently. |
| Linkup Search · pricing | sourced answers, structured output, deep search | Linkup packages retrieval and synthesis, but its general upstream index ownership is less transparent than independent-index providers. |
| Perplexity API · pricing | raw Search, Sonar, Agent API | Search request fees and model-token fees must be separated when comparing Perplexity with raw retrieval APIs. |
| Valyu Search · pricing | licensed finance, science, health, and legal data | Valyu is differentiated by specialist and proprietary datasets; its retrieval-based pricing is not directly comparable with per-request pricing. |
| Firecrawl Search · pricing | search, scrape, crawl, JavaScript, PDF | Firecrawl is primarily a content-acquisition and extraction layer and can be paired with a separate ranking provider. |
| Jina Reader | search-to-text and page-to-Markdown | Jina is a lightweight normalization layer; token volume and upstream ranking transparency should be tested before using it as the sole search backend. |
| OpenAI web search · pricing | model-native grounding and citations | Model-native search minimizes integration work but introduces model coupling and potentially multiple billable searches per turn. |
| Anthropic web search | server-side web search and fetch | Claude-native search and fetch are convenient inside Claude agent loops but require separate data-policy and portability review. |
| Gemini Grounding with Google Search · pricing | Google grounding, multilingual and local intent | A grounded prompt may generate multiple search queries, so billing should be measured at the solved-query level. |
| xAI pricing | Web Search and X Search | xAI is differentiated by first-class X retrieval; general-web ranking transparency remains limited. |
| Microsoft Grounding with Bing · Bing Search API retirement | Foundry grounding and legacy API migration | Grounding with Bing is platform-bound and is not a raw drop-in replacement for the retired Bing Search APIs. |
| Serper.dev · DataForSEO SERP API · SerpAPI | SERP proxies and aggregation | SERP list prices exclude page extraction, reranking, model input, retries, and citation work; compare cost per solved query. |
| Parallel benchmarks · You.com evaluation guidance | benchmark methodology | Vendor benchmarks often measure search plus reasoning under different tool-call and context budgets; use them for shortlisting, not final selection. |
Model Context Protocol
| Source | Topic | Usable conclusion |
|---|---|---|
| The 2026-07-28 Specification announcement | Release overview, SDK availability, ecosystem impact | MCP is moving from session-oriented bidirectional transport semantics toward stateless, cacheable, routable request/response infrastructure. |
| MCP 2026-07-28 specification | Authoritative protocol requirements | The normative specification defines stateless self-contained requests, per-request capability negotiation, optional extensions, and explicit consent and tool-safety responsibilities. |
| MCP 2026-07-28 key changes | Breaking changes, deprecations, migration details | The changelog is the implementation baseline for removing sessions, adopting MRTR and Tasks, adding gateway and cache metadata, hardening OAuth, and migrating deprecated features. |
Feishu MCP and Privileged Configuration
| Source | Topic | Usable conclusion |
|---|---|---|
| MCP overview · Local OpenAPI MCP overview · tool selection | local OpenAPI MCP, supported-tool boundary, -t allowlist | MCP is a tool-oriented interface to supported OpenAPI operations. It is suitable for a deliberately small agent tool surface, not an unrestricted administrative shell. |
| Developer Remote MCP · Personal Remote MCP | remote authentication, per-connection tool allowlist, personal-route lifecycle | Remote MCP is a separately controlled surface. Its documented developer tools are currently document-centred; personal MCP-token use is being superseded by Feishu CLI. |
| API permission introduction · application data permissions · administrator authorization | app scopes, data range, tenant approval, high-sensitive scope | Owner status does not bypass application scope, token identity, data permissions, or approval. High-sensitive scopes cannot be self-escalated through the administrator-authorization API. |
Testing / Browser Automation
| 来源 | 主题 | 可用结论 |
|---|---|---|
| Playwright Best Practices | user-visible behavior、test isolation、locators、web-first assertions、trace | E2E 应验证用户可见行为,保持测试隔离,优先语义 locator 和自动等待断言。 |
| Playwright homepage | Playwright Test、CLI、MCP | Playwright 同时覆盖确定性测试、agent 浏览器自动化和 MCP 场景,但长期质量门禁仍应落成测试代码。 |
| Testing Library Guiding Principles | user-centric testing | 测试越接近软件真实使用方式,越能提供信心。 |
| Static vs Unit vs Integration vs E2E Testing for Frontend Apps | Testing Trophy、confidence / cost tradeoff | 测试分层应围绕信心、速度和维护成本取舍,不是追求某一种测试形态。 |
| Chrome DevTools MCP for your AI agent | Chrome DevTools MCP、runtime debugging | Chrome DevTools MCP 给 coding agents 浏览器运行时观察和调试能力。 |
| Chrome DevTools for agents 1.0 | DevTools for agents、quality audits、emulation、memory leaks | Chrome DevTools for agents 适合运行时验证、审计、性能和真实浏览器问题诊断。 |
LangSmith / LangChain
| 来源 | 主题 | 可用结论 |
|---|---|---|
| AI Agent Observability: Tracing, Testing, and Improving Agents | traces、datasets、offline/online evals、production feedback | 生产 trace 是改进 agent 的燃料;失败样例应进入 dataset。 |
| LangSmith Evaluation | offline evaluation、online evaluation、datasets | 可连接本库 [[60 - 项目 Projects/LangSmith Agent Engineer Guide/06-evaluation-observability-playbook |
mattpocock/skills — primary sources
| Source | Topic | Usable conclusion |
|---|---|---|
| README at ed37663 | Positioning, installation, and user/model invocation | The skill-first goal is composable engineering discipline, not ownership of a complete development method. |
| Invocation design | User-invoked vs model-invoked behavior and Codex metadata | Invocation authority is the key split: explicit orchestration pays cognitive load, while automatic reuse pays context load. |
| Workflow source: ask-matt, setup, to-tickets, implement | Setup → grilling → spec/tickets → implement/TDD/review | Continuity relies on external artifacts such as docs, tracker state, specs, tickets, and ADRs rather than an internal state service. |
| Claude / Codex distribution ADR | skills.sh, Claude plugin, and Codex-plugin constraints | The upstream snapshot treats Claude as the managed-plugin path and defers a native Codex plugin. |
| v1.1.0 release · main comparison | Release versus main boundary | Pin behavior claims to a commit and distinguish formal releases from main’s development state. |
| Issue #558 | Codex / Claude dual-harness setup | The CLAUDE.md-first selection rule is a publicly tracked compatibility risk; it does not establish identical behavior in every client. |
Harness 选型 / 社区反馈
| 来源 | 类型 | 可用结论 |
|---|---|---|
| bmad-code-org/BMAD-METHOD | 官方仓库 | method-first:用 PRD、epics、stories、implementation、review 和 retrospective 组织完整研发流程。 |
| BMAD issue #446 | 历史 case study | v4.35.3 的使用者高度评价 elicitation,同时记录单体文档、agent isolation 和手工 handoff 等摩擦;不能直接外推到当前版本。 |
| BMAD issue #1332 | 已关闭 issue | 强制 review 最少发现问题数会产生 Goodhart 式副作用,说明确定性规则必须绑定正确指标。 |
| BMAD Reddit 使用反馈 | 社区样本 | Story file 对长周期和 brownfield 有价值,但 persona、状态和文档同步会带来真实 ceremony。 |
| mattpocock/skills | 官方仓库 | skill-first:把 grilling、TDD、code review、架构改进等能力按需组合。 |
| mattpocock/skills Discussion #214 | 社区样本 | 用户认可逐步追问对减少开发偏航的作用,也报告上手示范不足和长时间推理成本。 |
| Hacker News: Using AI to write better code more slowly | 社区讨论 | 多轮设计、拆小实现和 review 能提高理解与信心,但可能把速度收益转化为新的 review 工作。 |
| mindfold-ai/Trellis | 官方仓库 | repo-memory-first:用 specs、tasks、workspace memory、journals 和多平台 workflow 保持连续性。 |
| Trellis issue #415 | Open issue | 多 worktree 的 session 记录可能产生 merge conflict,持久状态会成为新的协作表面。 |
| Trellis issue #441 | 已关闭 issue | 无上限上下文注入曾造成 payload 膨胀风险,说明 repo memory 需要按需读取和预算。 |
| Trellis issue #402 | 已关闭 issue | 扁平 task 目录会隐藏父子关系,状态持久化仍需要良好的导航。 |
二手线索
- What is an AI agent harness? — 对 harness 作为 control plane 的解释清晰,可作概念补充。
- Harness Engineering — 适合连接旧的 AI Harness 101 参考。
- GitHub awesome list / Medium / Reddit 只作为发现来源,不要作为本包主证据。