Mails.ai with LangGraph

Every node checks a reply’s injection score before acting on it, and the graph’s state survives restarts in Postgres. A customer can take days to answer and nothing is lost.

Daichi OkudaML Engineer, Thrushwold

Install and wire it up

Install with pip install langgraph "langchain[mcp]" langchain-anthropic mailsai. The mailsai Python SDK (on PyPI) calls the REST API at https://api.mails.ai/v1 with Authorization: Bearer mk_live_…. Then wrap it as function tools and bind to your StateGraph nodes:

from langgraph.graph import StateGraph, MessagesState, START
from langgraph.prebuilt import ToolNode, tools_condition
from langchain.tools import tool
from langchain_anthropic import ChatAnthropic
from mailsai import agent
import os

# 1. Wire mails.ai
hello = agent("hello", api_key=os.environ["MAILS_API_KEY"])

@tool
def send_email(to: str, subject: str, body: str) -> dict:
    """Send an email from the hello agent."""
    result = hello.send(to=to, subject=subject, body=body)
    return {"send_id": result["id"], "thread_id": result["thread_id"]}

@tool
def list_recent_threads(limit: int = 20) -> list:
    """List recent inbound threads on the hello agent."""
    return hello.client.list_threads(agent_id="hello", limit=limit)["data"]

# 2. Bind to a model node
model = ChatAnthropic(model="claude-sonnet-5-5").bind_tools([send_email, list_recent_threads])

# 3. Wire into a StateGraph
graph = StateGraph(MessagesState)
graph.add_node("agent", lambda s: {"messages": [model.invoke(s["messages"])]})
graph.add_node("tools", ToolNode([send_email, list_recent_threads]))
graph.add_edge(START, "agent")
graph.add_conditional_edges("agent", tools_condition)
graph.add_edge("tools", "agent")
app = graph.compile()

See the graph send

Install the packages, run the graph, and its node calls send_email through the mailsai SDK.

Resume on a reply

Point a webhook at your FastAPI app. It checks the signature, skips a reply scored as an injection, and resumes the graph thread it belongs to.

@app.post("/webhooks/mails")
async def inbound(req: Request):
    body = await req.body()
    t, *sigs = req.headers["x-mails-signature"].split(",")
    msg = t[2:].encode() + b"." + body
    want = "v1=" + hmac.new(SECRET, msg, "sha256").hexdigest()
    ok = any(hmac.compare_digest(s, want) for s in sigs)
    if not ok or abs(time.time() - int(t[2:])) > 300:
        return Response(status_code=400)
    event = json.loads(body)  # reply.received
    if event["injection_score"] < 0.5:
        await resume(event)  # graph.ainvoke on its thread_id
    return Response(status_code=204)

“Replies land on our webhook, and the graph picks the thread back up where it waited, with the reply’s text and its injection score. A customer who answers a week later is no trouble.”

Kojo AmankwahSoftware Engineer, Mossgrebe

Guide

Function tools for orchestrated agent flows

Py SDK · 6 min read · docs.langchain.com

LangGraph is one of the most-downloaded Python orchestration frameworks for agent workflows — StateGraph for explicit state machines, conditional edges for branching logic, persistent checkpointing for long-running flows. Wire mails.ai as function tools and any node in your graph can send + read email as part of the workflow.

Why LangGraph + mails.ai

LangGraph’s explicit-state-machine model maps naturally onto agent-mail workflows where the state of the conversation matters across many steps:

  • Lifecycle notification pipelines. StateGraph nodes for detect_event → compose → send → wait_for_reply → classify_response → conditional branches based on intent. Persistent state so each user’s thread progresses independently.
  • Support escalation flows. Inbound typed event triggers a graph; the state machine walks through triage → auto-reply → (escalate-to-human if urgency > 0.8, with classify_inbound on) → resolve. The whole conversation lifecycle is graph state.
  • Lifecycle automations driven by an agent. Onboarding sequences where each step depends on what the user did or said in the previous email. State machine + persistent checkpointing means the user can take days between replies without state loss.

The two integration paths

Path A — @tool function wrappers (recommended). The install card pattern above. Idiomatic LangGraph code, one dependency (mailsai), full type hints for the state graph.

Path B — MCP server via langchain-mcp-adapters. If your team already uses MCP across other runtimes (Claude Code, Cursor, OpenAI Agents SDK) and wants consistency:

import os
from langchain_anthropic import ChatAnthropic
from langchain_mcp_adapters.client import MultiServerMCPClient
from langchain_core.messages import HumanMessage
from langchain.agents import create_agent

model = ChatAnthropic(model="claude-sonnet-5-5")
client = MultiServerMCPClient(
    {
        "mails": {
            "command": "npx",
            "args": ["-y", "@mailsai/mcp-server"],
            "env": {"MAILS_API_KEY": os.environ["MAILS_API_KEY"]},
            "transport": "stdio",
        },
    }
)
agent_runnable = create_agent(model, await client.get_tools())
result = await agent_runnable.ainvoke(
    {"messages": [HumanMessage(content="Email user@example.com a follow-up")]}
)

Path B trades a dependency for MCP-consistency across your stack. Path A is simpler for teams using LangGraph as their primary framework.

Inbound handling patterns

# Pattern: webhook -> StateGraph
from fastapi import FastAPI, Request
from langgraph.checkpoint.postgres.aio import AsyncPostgresSaver

app = FastAPI()

@app.post("/webhooks/mails-inbound")
async def handle_inbound(req: Request):
    event = await req.json()

    async with AsyncPostgresSaver.from_conn_string(POSTGRES_URI) as checkpointer:
        await checkpointer.setup()
        graph = your_support_graph.compile(checkpointer=checkpointer)

        # Inject the typed event into graph state
        config = {"configurable": {"thread_id": event["thread_id"]}}
        await graph.ainvoke(
            {
                "messages": [
                    {"role": "user", "content": f"Inbound from {event['data']['from']['address']}: {event['data']['extracted_text']}"},
                ],
                "current_event": event,  # the typed event: injection_score, and intent and entities with classification on
            },
            config=config,
        )
    return {"ok": True}

Common patterns

  • State-resumable email loops. The Postgres checkpointer persists conversation state. If your agent emails a user and waits 3 days for a reply, the graph state is preserved — the next inbound webhook resumes the graph from the waiting state.
  • Parallel fan-out for notifications. Use LangGraph’s parallel-edge patterns to dispatch individual 1:1 transactional notifications across many independent user threads. Pass a stable idempotency_key to hello.send (e.g. user id + step) so a retried node does not send twice.
  • Conditional branching on intent. The typed event’s intent field (turn on classify_inbound for the agent) becomes the branching condition for routing to different state graph paths (refund_request → refund-handling subgraph, schedule_demo → calendar-booking subgraph).

Security considerations

  • Injection guard before state injection. Webhook handler should hold the event for a person, before it reaches graph state, when it is quarantined, has no injection_score (it was not scanned) or scores 0.5 or more. Once it’s in state, the agent will reason against it.
  • Per-graph mails.ai key. Different StateGraphs (different agent identities) get different mails.ai keys. Key rotation per-graph without affecting other workflows.
  • Send budget enforcement at the graph layer. StateGraph state can track per-thread send count and conditionally short-circuit before exceeding a per-thread limit. Belt-and-suspenders to the per-agent send limits the platform already enforces, and to its per-plan send caps (/docs/limits).

Compare against the Anthropic SDK setup for a more direct (non-orchestration) integration, or Pydantic AI for typed-input agent definitions.

Read next: Mails.ai with Pydantic AI and Mails.ai with the Anthropic SDK.

Where to go next

The quickstart, both SDKs and the full API reference.

  • Quickstart

    From an API key to a sent message and its reply event, over plain REST.

    Read the quickstart
  • SDKs

    @mailsai/sdk for TypeScript and mailsai for Python, thin wrappers over the same API.

    Read the SDK docs
  • API reference

    Every endpoint, with the exact request and response fields.

    Check the reference

Questions developers ask after wiring this up.

Why function tools instead of MCP for LangGraph?
LangGraph doesn’t have first-class MCP-client primitives in its core API today. LangChain’s langchain.mcp.MCPAdapter (beta) loads MCP tools into LangGraph, but it adds a dependency layer and the @tool function approach is more idiomatic LangGraph code. If your team prefers MCP for consistency with other runtimes, the adapter works — just add `pip install "langchain[mcp]"` and wrap the mails MCP server.
How do I handle inbound webhooks in a LangGraph workflow?
Two patterns. (1) Polling: a scheduled LangGraph run that calls hello.list_replies(since=…), processes new typed events, and updates state. Good for low-volume async flows. (2) Webhook-driven: a separate FastAPI endpoint receives the webhook from mails.ai, persists the event to your StateGraph’s checkpointer, and triggers the relevant graph from the right state. The second pattern matches how LangGraph customers usually wire external triggers.
Does this work with LangGraph’s persistent state / checkpointing?
Yes. mails.send results (id, thread_id, classifier_score) can be stored in your StateGraph’s state — useful for resumable flows where the graph picks up after an email is sent and waits for a reply. The checkpointer (PostgresSaver, SqliteSaver, etc.) persists the full state graph including any mails.* tool call results.
Can I run mails.* tools in parallel LangGraph nodes?
Yes. mails.send is idempotent when you pass the same idempotency_key on every retry (kept 24 hours); derive it from the user and step. Parallel fan-out (e.g., dispatching individual 1:1 transactional notifications across many independent user threads via parallel branches in StateGraph) works without duplicate-send risk. Per-key, per-agent and per-plan send limits apply — see /docs/limits.

“Replies come back as events with an injection score already on them. We deleted a whole layer of parsing code the week we switched, and we gate on the quarantine flag, so our agent never sees the ones that look like attacks.”

Tomás VargaStaff Engineer, Cinderjay

“Moving our agent onto our own domain was a few DNS records at the registrar. No nameserver move, and our existing mail kept working. Replies to the agent still come back to its inbox, threaded with the message they answer.”

Ishani VaidyaCTO, Sedgequay

“The 422 on cold outreach is the feature I didn’t know I wanted. An agent can’t talk itself into emailing strangers.”

Marcus FeldFounder, Marrowkite

“Our tests send to the test address and wait for the real reply, so the whole loop is covered before a customer ever writes in. It answers in about a second, which keeps the suite fast.”

Ana Lucía RíosEngineering Lead, Gorsefinch

“Adding the MCP server was one JSON block. Claude Code could send and read its own inbox a minute later.”

Jonah AbramsDeveloper, Wickerjay

“A reputation score per agent tells us exactly which one needs attention, instead of one number for the whole account. It comes from each agent’s own replies, bounces and complaints, and we read it from the API.”

Mei Lin ZhouOperations, Larchwhistle

“Sends and replies have separate allowances, so a busy inbox never eats our sending quota. Pricing was the easy part.”

Kwabena AdomakoFounder, Oxbowlark

“Half our agents are LangGraph in Python and half are Node. Both SDKs make the same calls, so the team doesn’t have to think about it.”

Sofia BrandtEngineer, Moss & Ladder

“Signed webhooks, retries with backoff and an event id to dedupe on. The boring plumbing, done properly.”

Nikhil BhonsleBackend Lead, Thimblecrow

“We signed up, made a key and sent our first message without talking to anyone, and the free tier never asked for a card. That’s how infrastructure should feel.”

Frieda WesselFounder, Ploverwick Studio

Give your first agent an inbox

Free covers 3,000 emails and 3,000 inbound replies a month, with no card. Upgrade when your agents get busy.

Get your API key