1. 依赖
FastAPI 与 LangChain/LangGraph 项目通常会遇到三类测试:
| 类型 | 关注点 | 是否访问外部服务 |
|---|---|---|
| 单元测试 | Settings、Service、节点和工具函数 | 不访问 |
| 接口测试 | 路由、校验、依赖、状态码和响应 | 通常不访问 |
| 集成测试 | 数据库、模型供应商、checkpointer | 根据环境决定 |
测试工具属于开发依赖:
1uv add --group test pytest pytest-asyncio coverage httpx asgi-lifespan
对应 TOML:
1[dependency-groups]2test = [3"asgi-lifespan",4"coverage",5"httpx",6"pytest",7"pytest-asyncio",8]
这里省略版本约束,只是为了突出依赖归属。实际执行 uv add 时,uv 会根据项目策略更新 pyproject.toml 和 uv.lock,CI 再按照锁文件安装已经确认的版本。
FastAPI 的同步 TestClient 能覆盖很多接口,但 LangChain 模型、异步数据库和 LangGraph checkpointer 经常使用异步 API。pytest-asyncio 负责运行异步测试,HTTPX 可以直接请求 ASGI 应用,asgi-lifespan 则负责在异步接口测试前后触发应用的启动与关闭流程。
生产镜像通常不需要安装这些测试工具。CI 可以保留项目运行依赖,只额外安装 test 组:
1uv sync --locked --no-default-groups --group test2uv run --no-sync pytest -m "not integration"
--locked 会在锁文件与 pyproject.toml 不一致时让 CI 失败。--no-default-groups 排除默认开发组,--group test 再加入测试依赖,同时保留项目本身及其运行依赖。后面的 --no-sync 表示直接使用刚刚同步好的环境,避免 uv run 再次按默认依赖组调整环境。
不要把这里改成 --only-group test。only 会连项目本身和运行依赖一起排除,测试导入 FastAPI、LangChain 或 app 包时就可能失败。它只适合明确不需要项目代码的独立工具环境。
2. pytest
pytest 可以从 [tool.pytest.ini_options] 读取配置:
1[tool.pytest.ini_options]2minversion = "8.4"3testpaths = ["tests"]4python_files = ["test_*.py", "*_test.py"]5python_classes = ["Test*"]6python_functions = ["test_*"]7addopts = "-ra --strict-config --strict-markers"
常用字段含义如下:
| 字段 | 含义 |
|---|---|
minversion | 运行测试所需最低 pytest 版本 |
testpaths | 未传路径时默认搜索目录 |
python_files | 哪些文件被识别为测试文件 |
python_classes | 哪些类被识别为测试类 |
python_functions | 哪些函数被识别为测试函数 |
addopts | 每次 pytest 默认附加的命令行选项 |
-ra 会在结尾显示跳过、预期失败等非通过结果摘要。--strict-config 会把未知配置视为错误,--strict-markers 会阻止拼错 marker 后测试被错误分类。
不要在 addopts 中默认加入 -x。-x 遇到第一个失败就停止,适合本地快速排查,但 CI 通常需要看到完整失败列表。
pytest 9 开始支持原生 TOML 表 [tool.pytest],但 [tool.pytest.ini_options] 仍然受支持。本文保留后者,是为了兼容 minversion = "8.4";同一个 pyproject.toml 中不要同时配置这两个表。
测试目录应与应用模块对应:
目录组织不是 pytest 强制要求,但它能帮助团队选择测试范围并区分是否需要外部资源。
不确定 pytest 找到了哪些测试时,可以先只查看收集结果,也可以按目录、文件或节点 ID 运行:
1uv run pytest --collect-only2uv run pytest tests/unit3uv run pytest tests/api/test_health.py::test_health
测试函数参数中的 test_settings、app、model 等名称通常是 fixture,不是忘记定义的普通变量。pytest 会按参数名寻找当前文件、插件或 conftest.py 中同名的 fixture,再把返回值注入测试函数。如果名称不存在,pytest 会报告 fixture not found;可以运行 uv run pytest --fixtures 查看当前可用的 fixture。
3. 标记
markers 用于给测试分类:
1[tool.pytest.ini_options]2markers = [3"unit: 不访问网络和数据库的快速测试",4"api: FastAPI 接口与依赖测试",5"integration: 访问真实数据库或外部服务",6"slow: 运行时间较长的测试",7]
测试中使用 marker:
1import pytest2from langchain_core.language_models import BaseChatModel345@pytest.mark.integration6@pytest.mark.slow7async def test_real_model_connection(model: BaseChatModel) -> None:8response = await model.ainvoke("只回复 OK")9assert response.text.strip()
这里的 model 应由集成测试专用 fixture 创建。它与后面的假模型不同,会连接真实供应商,因此只应在明确选择集成测试并提供测试密钥时运行。
选择测试:
1uv run pytest -m unit2uv run pytest -m "api and not slow"3uv run pytest -m "not integration"
marker 只是分类,不会自动阻止网络请求。单元测试是否真的离线,仍取决于依赖替换和实现。不要给一个会调用真实 DeepSeek 的测试标记为 unit 后就认为它变快了。
跳过真实集成测试时应写明条件:
01import os0203import pytest04from langchain_core.language_models import BaseChatModel0506pytestmark = pytest.mark.integration070809@pytest.mark.skipif(10not os.getenv("DEEPSEEK_API_KEY"),11reason="没有配置 DeepSeek API Key",12)13async def test_real_model_connection(model: BaseChatModel) -> None:14response = await model.ainvoke("只回复 OK")15assert response.text.strip()
CI 可以设置独立任务,在受控环境中使用权限受限的测试密钥运行集成测试;普通 Pull Request(PR)只运行离线测试。不要把生产密钥交给来自外部贡献者的工作流。
4. 异步
pytest-asyncio 可以在 pyproject.toml 中配置异步模式:
1[tool.pytest.ini_options]2asyncio_mode = "auto"3asyncio_default_fixture_loop_scope = "function"
asyncio_mode = "auto" 会自动接管异步测试函数和异步 fixture,因此示例只写 @pytest.mark.api 也可以运行。如果改为 strict 模式,异步测试需要显式添加 @pytest.mark.asyncio,异步 fixture 则应使用 @pytest_asyncio.fixture。两种模式没有绝对优劣,关键是项目保持一致,不要只改配置却忘记调整装饰器。
fixture 的事件循环作用域会影响异步资源是否跨测试复用。函数级作用域最容易理解,每个测试结束后都会释放自己的循环;数据库连接池等昂贵资源如果需要跨测试复用,可以显式扩大 loop_scope。事件循环的作用域不能小于异步 fixture 自身的缓存作用域,而且资源必须能够可靠清理,否则测试之间会互相污染。
FastAPI 异步接口可以使用 HTTPX:
01import pytest02from asgi_lifespan import LifespanManager03from fastapi import FastAPI04from httpx import ASGITransport, AsyncClient050607@pytest.mark.api08async def test_health(app: FastAPI) -> None:09async with LifespanManager(app) as manager:10transport = ASGITransport(app=manager.app)1112async with AsyncClient(13transport=transport,14base_url="http://testserver",15) as client:16response = await client.get("/api/v1/health")1718assert response.status_code == 200
HTTPX 的 ASGITransport 不会自动触发 ASGI lifespan。这里用 LifespanManager 显式执行启动和关闭,并把 manager.app 交给 transport,这样生命周期状态也能传入请求。前一篇文章中的 app fixture 注入了空 lifespan,所以普通接口测试不会连接真实模型和数据库;测试资源型路由时,应改为注入假资源的测试 lifespan。
测试异步代码时不要使用真实 sleep() 等待几秒。定时、重试和超时逻辑应允许注入时钟或缩短测试参数,否则测试套件会越来越慢。
5. 配置测试
配置系统本身也需要测试。可以使用 pytest 的 tmp_path 创建临时 TOML:
01from pathlib import Path0203import pytest0405from app.core.config_loader import ConfigError, load_config0607VALID_CONFIG = """08[app]09name = "FastAPI LangChain Test"10environment = "test"11log_level = "WARNING"1213[server]14host = "127.0.0.1"15port = 80001617[server.cors]18allow_origins = ["http://testserver"]19allow_methods = ["GET", "POST"]2021[model]22provider = "fake"23name = "test-model"24base_url = "https://example.test"25temperature = {temperature}26timeout_seconds = 527max_retries = 02829[langgraph]30recursion_limit = 1031checkpoint_backend = "memory"3233[database]34pool_size = 135pool_timeout_seconds = 536"""373839def write_config(tmp_path: Path, *, temperature: float = 0.1) -> Path:40config_path = tmp_path / "config.toml"41config_path.write_text(42VALID_CONFIG.format(temperature=temperature),43encoding="utf-8",44)45return config_path464748def test_load_model_config(tmp_path: Path) -> None:49config = load_config(write_config(tmp_path))5051assert config.model.name == "test-model"52assert config.model.temperature == 0.1535455def test_reject_invalid_temperature(tmp_path: Path) -> None:56with pytest.raises(ConfigError, match=r"model\.temperature"):57load_config(write_config(tmp_path, temperature=9))
AppConfig 要求五个配置分组都存在,所以测试基准文件也要提供一份最小但完整的合法配置。只写 [model] 会因为缺少其他分组而失败,无法证明错误确实来自 temperature。这里的加载器会把底层 ValidationError 转换为项目自己的 ConfigError,测试也应断言对外暴露的异常类型和字段路径。
tmp_path 为每个测试提供独立临时目录,测试结束后由 pytest 清理。这样既能验证真实文件加载,又不会改写仓库中的 config.toml。
环境变量覆盖使用 monkeypatch:
01import pytest0203from app.core.settings import get_settings040506def test_environment_overrides_toml(07monkeypatch: pytest.MonkeyPatch,08) -> None:09monkeypatch.setenv("MODEL__NAME", "environment-model")10get_settings.cache_clear()1112try:13settings = get_settings()14assert settings.model.name == "environment-model"15finally:16get_settings.cache_clear()
monkeypatch 会在测试结束后恢复环境变量,但不会自动清理 get_settings() 的 lru_cache。如果不在前后清除缓存,这个测试可能读到其他用例创建的 Settings,或者把自己的结果泄漏给后续测试。运行这类测试时,项目的测试配置仍需提供其他必填字段。
测试至少覆盖:
- TOML 默认值能够读取;
- 环境变量拥有预期优先级;
- 真正调用模型或数据库的工厂在缺少所需密钥时明确失败;
- 数字、布尔值和数组能正确转换;
- 非法环境名、模型参数和数据库设置被拒绝;
- 错误日志不会暴露密钥等敏感值。
6. 覆盖率与 AI
Coverage.py 可以从 [tool.coverage.*] 读取 TOML:
01[tool.coverage.run]02branch = true03source = ["app"]0405[tool.coverage.report]06show_missing = true07skip_covered = true08fail_under = 8509exclude_also = [10"if TYPE_CHECKING:",11"raise NotImplementedError",12]1314[tool.coverage.html]15directory = "coverage_html"
字段含义:
| 字段 | 含义 |
|---|---|
branch | 除了行,还检查 if/else 等分支是否覆盖 |
source | 需要测量的项目代码范围 |
show_missing | 报告中显示未覆盖行号 |
skip_covered | 报告中隐藏完全覆盖文件 |
fail_under | 总覆盖率低于阈值时失败 |
exclude_also | 从覆盖统计中排除特定代码模式 |
执行:
1uv run coverage run -m pytest -m "not integration"2uv run coverage report3uv run coverage html
行覆盖率高不代表测试有效。下面的测试虽然执行了 Service,却没有验证任何业务结果:
1from app.services.chat import ChatService234async def test_chat_service(chat_service: ChatService) -> None:5await chat_service.answer("hello")
更重要的是验证可观察行为、错误分支和边界条件。fail_under 应用于阻止覆盖率无意下降,而不是逼迫团队为每一行写没有意义的测试。
这里的 85 只是示例阈值。已有项目更适合先记录当前基线,再逐步提高;新项目则可以从一开始约定目标。由于 source = ["app"] 已经只测量应用代码,不需要再用 omit 排除 tests/。
不要为了提高数字把核心错误处理和模型解析逻辑加入 omit。排除项应局限于测试无法或没有必要执行的声明性代码。
AI 测试
真实模型输出具有成本、延迟和不确定性,不能作为普通单元测试的默认依赖。Service 应依赖 LangChain 的模型抽象,测试传入假模型:
01from langchain_core.language_models.fake_chat_models import (02FakeMessagesListChatModel,03)04from langchain_core.messages import AIMessage050607def create_fake_model() -> FakeMessagesListChatModel:08return FakeMessagesListChatModel(09responses=[AIMessage(content="固定测试回答")],10)
FakeMessagesListChatModel 本身就是 BaseChatModel 的测试实现,支持同步和异步调用,比返回任意对象或手写一个不完整的假类更接近真实接口。它只返回预设消息,不会访问网络。
对 LangGraph,节点逻辑、路由条件和状态更新可以分别测试;完整图测试再使用内存 checkpointer 和固定 thread_id。生产 PostgreSQL checkpointer 留给集成测试。
AI 配置测试应验证参数是否正确传给工厂,而不是只验证 TOML 能解析:
01from unittest.mock import patch0203from pydantic import SecretStr0405from app.agent.model import create_model06from app.core.settings import Settings070809def test_model_factory_uses_settings(test_settings: Settings) -> None:10settings = test_settings.model_copy(11update={12"model": test_settings.model.model_copy(13update={"api_key": SecretStr("test-api-key")},14),15},16)1718with patch("app.agent.model.ChatOpenAI") as chat_openai:19create_model(settings)2021call = chat_openai.call_args22assert call is not None23assert call.kwargs["model"] == "fake-model"24assert call.kwargs["api_key"] == "test-api-key"25assert call.kwargs["timeout"] == 526assert call.kwargs["max_retries"] == 0
这里复制 test_settings 并加入专用假密钥,不会修改其他测试共享的 fixture。patch() 的目标是模型工厂实际查找 ChatOpenAI 的位置,也就是 app.agent.model.ChatOpenAI,而不是这个类最初定义的第三方模块。测试只验证构造参数,不会发出模型请求。
追踪默认在测试中关闭,避免把测试提示词和假用户数据发送到 LangSmith。专门验证 tracing 的集成测试可以在独立环境显式开启。
测试失败时要能复现。涉及随机温度、当前时间、重试抖动和并发顺序时,尽量注入可控制依赖,避免用不断重跑来掩盖偶发失败。