03 来源索引

返回 AI Harness

来源优先级

  1. OpenAI / Anthropic 官方工程博客、开发者文档、官方 guide。
  2. Google / Microsoft 等大厂官方云平台和工程博客。
  3. LangSmith / Langfuse / Braintrust 等 eval/observability 工具官方资料。
  4. 二手博客、awesome list、社区帖只做线索,不做主证据。

OpenAI

来源主题可用结论
A practical guide to building agentsagent foundations、orchestration、guardrails可靠 agent 需要模型、工具、结构化指令、护栏和渐进式部署。
New tools for building agentsResponses API、Agents SDK、tools、handoffs、guardrails、tracingOpenAI 把 tools、handoffs、guardrails、tracing 作为 agent 应用 primitives。
Safety in building agentsinput guardrails、PII、jailbreak、trace graders、tool approvals对不可信输入用结构化字段、隔离、tool approval 和 human approval 降低风险。
Evaluate agent workflowstraces、graders、datasets、eval runs从 trace 调试行为,再沉淀 dataset 和 eval runs 做可重复评估。
How to implement LLM guardrailsguardrail cookbook可作为实现输入/输出 guardrails 的代码参考。
Building Governed AI Agentsgoverned agents、scaffolding作为 governance 和 tool policy 的补充参考;优先级低于官方 guide。

Anthropic

来源主题可用结论
Trustworthy agents in practicemodel / harness / tools / environmentAnthropic 明确把 harness 定义为模型运行其上的 instructions 和 guardrails。
Effective context engineering for AI agentscontext、tools、examples、message history、runtime retrievalcontext engineering 是管理完整 context state,不是只改 prompt。
Writing effective tools for agentstool design、high-signal output、eval-driven iteration工具要清晰、少重叠、返回高信号结果,并通过 eval/transcript 迭代。
Building Effective AI Agentsworkflow patterns、simplicity、transparency、ACI简单、透明、良好的 agent-computer interface 是可靠 agent 的基础。
Demystifying evals for AI agentsagent harness、eval suite、evaluation harness、outcome评估 agent 时是在评估 model + harness;要看 trajectory 和 final outcome。
Harness design for long-running application developmentlong-running agents、context reset、handoff、self-evaluation长任务需要 context reset、structured handoff 和独立 evaluator。
How we built our multi-agent research systemmulti-agent、tool design、context/memory可作为复杂 agent 的旁证,尤其是工具描述、context 管理和子代理 handoff。

Google

来源主题可用结论
A dev’s guide to production-ready AI agentsproduction agents、trajectory eval、sandbox/canary/prodagent 评估要看完整决策路径,部署要分 sandbox、canary、production。
Startup technical guide: AI agentsADK、observability、evaluation、context managementADK 把 context management、evaluation、observability、containerization、multi-agent composition 放进生产能力。
Agent Factory Recap: Securing AI Agents in Productionnetwork isolation、logging、tool safeguards安全 harness 需要网络隔离、失败日志、工具保护和审计线索。

Microsoft

来源主题可用结论
Agent Factory: Top 5 agent observability best practices for reliable AIAzure AI Foundry、evals、red teaming、CI/CD、dashboardsobservability 要覆盖开发、CI/CD 和生产,并纳入 safety/governance。
Agent evaluators for generative AIagent evaluators、workflow quality生产级 agent 需要评估每一步 workflow,不只是 final output。
Observability for AI Systemsmulti-turn trace、evaluation、governanceAI observability 的关联范围应覆盖完整 conversation / memory lifecycle。

Agent Search and Web Retrieval

SourceTopicUsable conclusion
Parallel Search · pricingindependent index, excerpts, Search, Task, monitoringParallel is designed for agent-native research; Search and high-depth Task processors are separate cost and latency classes.
Exa Search API · pricingsemantic search, specialized indexes, structured outputExa is strongest as a semantic discovery layer for entities, research, and code rather than a traditional consumer SERP.
Tavily API · creditssearch, extract, crawl, map, researchTavily offers the broadest integrated retrieval workflow, but endpoint depth and credit consumption must be included in cost evaluation.
Brave Search API · pricingindependent index, web verticals, LLM ContextBrave is a strong independent raw-search provider and fallback when avoiding correlated Google/Bing dependencies matters.
You.com Search · billingsearch, contents, managed researchSearch, page contents, and multi-tier research are separate products and should be benchmarked and budgeted independently.
Linkup Search · pricingsourced answers, structured output, deep searchLinkup packages retrieval and synthesis, but its general upstream index ownership is less transparent than independent-index providers.
Perplexity API · pricingraw Search, Sonar, Agent APISearch request fees and model-token fees must be separated when comparing Perplexity with raw retrieval APIs.
Valyu Search · pricinglicensed finance, science, health, and legal dataValyu is differentiated by specialist and proprietary datasets; its retrieval-based pricing is not directly comparable with per-request pricing.
Firecrawl Search · pricingsearch, scrape, crawl, JavaScript, PDFFirecrawl is primarily a content-acquisition and extraction layer and can be paired with a separate ranking provider.
Jina Readersearch-to-text and page-to-MarkdownJina is a lightweight normalization layer; token volume and upstream ranking transparency should be tested before using it as the sole search backend.
OpenAI web search · pricingmodel-native grounding and citationsModel-native search minimizes integration work but introduces model coupling and potentially multiple billable searches per turn.
Anthropic web searchserver-side web search and fetchClaude-native search and fetch are convenient inside Claude agent loops but require separate data-policy and portability review.
Gemini Grounding with Google Search · pricingGoogle grounding, multilingual and local intentA grounded prompt may generate multiple search queries, so billing should be measured at the solved-query level.
xAI pricingWeb Search and X SearchxAI is differentiated by first-class X retrieval; general-web ranking transparency remains limited.
Microsoft Grounding with Bing · Bing Search API retirementFoundry grounding and legacy API migrationGrounding with Bing is platform-bound and is not a raw drop-in replacement for the retired Bing Search APIs.
Serper.dev · DataForSEO SERP API · SerpAPISERP proxies and aggregationSERP list prices exclude page extraction, reranking, model input, retries, and citation work; compare cost per solved query.
Parallel benchmarks · You.com evaluation guidancebenchmark methodologyVendor benchmarks often measure search plus reasoning under different tool-call and context budgets; use them for shortlisting, not final selection.

Model Context Protocol

SourceTopicUsable conclusion
The 2026-07-28 Specification announcementRelease overview, SDK availability, ecosystem impactMCP is moving from session-oriented bidirectional transport semantics toward stateless, cacheable, routable request/response infrastructure.
MCP 2026-07-28 specificationAuthoritative protocol requirementsThe normative specification defines stateless self-contained requests, per-request capability negotiation, optional extensions, and explicit consent and tool-safety responsibilities.
MCP 2026-07-28 key changesBreaking changes, deprecations, migration detailsThe changelog is the implementation baseline for removing sessions, adopting MRTR and Tasks, adding gateway and cache metadata, hardening OAuth, and migrating deprecated features.

Feishu MCP and Privileged Configuration

SourceTopicUsable conclusion
MCP overview · Local OpenAPI MCP overview · tool selectionlocal OpenAPI MCP, supported-tool boundary, -t allowlistMCP is a tool-oriented interface to supported OpenAPI operations. It is suitable for a deliberately small agent tool surface, not an unrestricted administrative shell.
Developer Remote MCP · Personal Remote MCPremote authentication, per-connection tool allowlist, personal-route lifecycleRemote MCP is a separately controlled surface. Its documented developer tools are currently document-centred; personal MCP-token use is being superseded by Feishu CLI.
API permission introduction · application data permissions · administrator authorizationapp scopes, data range, tenant approval, high-sensitive scopeOwner status does not bypass application scope, token identity, data permissions, or approval. High-sensitive scopes cannot be self-escalated through the administrator-authorization API.

Testing / Browser Automation

来源主题可用结论
Playwright Best Practicesuser-visible behavior、test isolation、locators、web-first assertions、traceE2E 应验证用户可见行为,保持测试隔离,优先语义 locator 和自动等待断言。
Playwright homepagePlaywright Test、CLI、MCPPlaywright 同时覆盖确定性测试、agent 浏览器自动化和 MCP 场景,但长期质量门禁仍应落成测试代码。
Testing Library Guiding Principlesuser-centric testing测试越接近软件真实使用方式,越能提供信心。
Static vs Unit vs Integration vs E2E Testing for Frontend AppsTesting Trophy、confidence / cost tradeoff测试分层应围绕信心、速度和维护成本取舍,不是追求某一种测试形态。
Chrome DevTools MCP for your AI agentChrome DevTools MCP、runtime debuggingChrome DevTools MCP 给 coding agents 浏览器运行时观察和调试能力。
Chrome DevTools for agents 1.0DevTools for agents、quality audits、emulation、memory leaksChrome DevTools for agents 适合运行时验证、审计、性能和真实浏览器问题诊断。

LangSmith / LangChain

来源主题可用结论
AI Agent Observability: Tracing, Testing, and Improving Agentstraces、datasets、offline/online evals、production feedback生产 trace 是改进 agent 的燃料;失败样例应进入 dataset。
LangSmith Evaluationoffline evaluation、online evaluation、datasets可连接本库 [[60 - 项目 Projects/LangSmith Agent Engineer Guide/06-evaluation-observability-playbook

mattpocock/skills — primary sources

SourceTopicUsable conclusion
README at ed37663Positioning, installation, and user/model invocationThe skill-first goal is composable engineering discipline, not ownership of a complete development method.
Invocation designUser-invoked vs model-invoked behavior and Codex metadataInvocation authority is the key split: explicit orchestration pays cognitive load, while automatic reuse pays context load.
Workflow source: ask-matt, setup, to-tickets, implementSetup → grilling → spec/tickets → implement/TDD/reviewContinuity relies on external artifacts such as docs, tracker state, specs, tickets, and ADRs rather than an internal state service.
Claude / Codex distribution ADRskills.sh, Claude plugin, and Codex-plugin constraintsThe upstream snapshot treats Claude as the managed-plugin path and defers a native Codex plugin.
v1.1.0 release · main comparisonRelease versus main boundaryPin behavior claims to a commit and distinguish formal releases from main’s development state.
Issue #558Codex / Claude dual-harness setupThe CLAUDE.md-first selection rule is a publicly tracked compatibility risk; it does not establish identical behavior in every client.

Harness 选型 / 社区反馈

来源类型可用结论
bmad-code-org/BMAD-METHOD官方仓库method-first:用 PRD、epics、stories、implementation、review 和 retrospective 组织完整研发流程。
BMAD issue #446历史 case studyv4.35.3 的使用者高度评价 elicitation,同时记录单体文档、agent isolation 和手工 handoff 等摩擦;不能直接外推到当前版本。
BMAD issue #1332已关闭 issue强制 review 最少发现问题数会产生 Goodhart 式副作用,说明确定性规则必须绑定正确指标。
BMAD Reddit 使用反馈社区样本Story file 对长周期和 brownfield 有价值,但 persona、状态和文档同步会带来真实 ceremony。
mattpocock/skills官方仓库skill-first:把 grilling、TDD、code review、架构改进等能力按需组合。
mattpocock/skills Discussion #214社区样本用户认可逐步追问对减少开发偏航的作用,也报告上手示范不足和长时间推理成本。
Hacker News: Using AI to write better code more slowly社区讨论多轮设计、拆小实现和 review 能提高理解与信心,但可能把速度收益转化为新的 review 工作。
mindfold-ai/Trellis官方仓库repo-memory-first:用 specs、tasks、workspace memory、journals 和多平台 workflow 保持连续性。
Trellis issue #415Open issue多 worktree 的 session 记录可能产生 merge conflict,持久状态会成为新的协作表面。
Trellis issue #441已关闭 issue无上限上下文注入曾造成 payload 膨胀风险,说明 repo memory 需要按需读取和预算。
Trellis issue #402已关闭 issue扁平 task 目录会隐藏父子关系,状态持久化仍需要良好的导航。

二手线索