Runtime lines

LangChain and LangGraph, the essential toolkit

LangChain gives you standard parts for LLM apps: chat models, messages, prompts, tools and retrievers, which all snap together as Runnables. LangGraph runs them as a graph with state, loops, memory and human approval. Step through four short programs, then read the parts, the patterns and the questions interviewers ask.

Every program on this page really runs, without an API key. The model is a fake chat model from langchain_core.language_models that replays scripted replies, so every printed output is real. In your app you write model = init_chat_model("provider:model-name") instead (with the provider’s package installed, such as langchain-openai or langchain-anthropic), and nothing else in the code changes.

A chain with LCEL

chain.pylangchain 1.4, langgraph 1.2
    Output
      The route the lit station is running
      edgeconditional edge
      Value between the steps

        The building blocks

        Everything in LangChain implements one interface, Runnable. Learn its methods once and you can call, compose, stream and test every part the same way: prompts, models, parsers, retrievers, tools, whole chains and compiled graphs.

        One interface for every Runnable

        runnables.py pick a call to see its output
          Output

            Chat models and messages

            init_chat_model("provider:model-name") returns a chat model for any supported provider with the same API. A model takes a list of messages and returns an AIMessage.

            • SystemMessage: instructions for the model.
            • HumanMessage: what the user says.
            • AIMessage: the reply. It can carry tool_calls and usage_metadata (token counts).
            • ToolMessage: a tool’s result, linked to its call by tool_call_id.

            In 1.x, message.text is a property that gives the plain text, and message.content_blocks gives typed blocks (text, reasoning, tool calls, images) in the same shape for every provider. Streaming yields AIMessageChunks, which you can add together with +.

            Prompt templates

            ChatPromptTemplate turns a dict into a list of messages. MessagesPlaceholder inserts a whole list, usually the chat history.

            from langchain_core.messages import AIMessage, HumanMessage
            from langchain_core.prompts import ChatPromptTemplate, MessagesPlaceholder
            
            prompt = ChatPromptTemplate.from_messages(
                [
                    ("system", "You are a {role}."),
                    MessagesPlaceholder("history"),
                    ("human", "{question}"),
                ]
            )
            value = prompt.invoke(
                {
                    "role": "tutor",
                    "history": [HumanMessage("Hi"), AIMessage("Hello!")],
                    "question": "What is LCEL?",
                }
            )
            print([(m.type, m.content) for m in value.to_messages()])
            [('system', 'You are a tutor.'), ('human', 'Hi'), ('ai', 'Hello!'), ('human', 'What is LCEL?')]

            Structured output and parsers

            model.with_structured_output(Schema) returns a Runnable that gives you a validated object instead of a message. It uses the provider’s native JSON mode or tool calling under the hood. Pass include_raw=True to also get the raw message and any parsing error.

            from langchain_core.language_models import GenericFakeChatModel
            from langchain_core.messages import AIMessage
            from pydantic import BaseModel
            
            
            class FakeChatModel(GenericFakeChatModel):
                def bind_tools(self, tools, **kwargs):
                    return self
            
            
            class Movie(BaseModel):
                """A movie mentioned in the text."""
            
                title: str
                year: int
            
            
            call = {"name": "Movie", "args": {"title": "Inception", "year": 2010}, "id": "c1"}
            model = FakeChatModel(messages=iter([AIMessage("", tool_calls=[call])]))
            
            extractor = model.with_structured_output(Movie)
            movie = extractor.invoke("Nolan's 2010 film about dreams")
            print(repr(movie))
            Movie(title='Inception', year=2010)

            Output parsers do the same job inside a chain: StrOutputParser for text, JsonOutputParser and PydanticOutputParser for models without native structured output.

            Tools

            @tool turns a typed function into a tool: the name comes from the function, the description from the docstring and the argument schema from the type hints (or a Pydantic args_schema). The model sees only those three, so write them for the model.

            • model.bind_tools([...]) lets the model decide to call them.
            • tool.invoke(tool_call) returns a ToolMessage with the right tool_call_id.
            • ToolNode runs all the calls in the last AIMessage in parallel. By default it sends invalid-argument errors back to the model and re-raises other exceptions; handle_tool_errors=True returns every error as a message.

            RAG: documents, splitters, embeddings, retrievers

            • Loaders read sources into Documents (page_content plus metadata).
            • Text splitters (package langchain-text-splitters) cut them into chunks. RecursiveCharacterTextSplitter tries paragraphs, then lines, then words; chunk_size caps each chunk and chunk_overlap repeats a little text between neighbours so an idea isn’t cut in half.
            • Embeddings turn text into vectors (init_embeddings("provider:model-name")); a vector store keeps them and finds the nearest ones.
            • store.as_retriever() gives a retriever: a Runnable from a query string to a list of documents, so it drops straight into a chain.
            from langchain_core.documents import Document
            from langchain_core.embeddings import DeterministicFakeEmbedding
            from langchain_core.language_models import FakeListChatModel
            from langchain_core.output_parsers import StrOutputParser
            from langchain_core.prompts import ChatPromptTemplate
            from langchain_core.runnables import RunnablePassthrough
            from langchain_core.vectorstores import InMemoryVectorStore
            
            docs = [
                Document("LangGraph adds state, cycles and checkpoints.", metadata={"src": "a"}),
                Document("LCEL joins Runnables with the | operator.", metadata={"src": "b"}),
            ]
            store = InMemoryVectorStore.from_documents(docs, DeterministicFakeEmbedding(size=64))
            retriever = store.as_retriever(search_kwargs={"k": 1})
            
            prompt = ChatPromptTemplate.from_template(
                "Answer from the context only.\nContext: {context}\nQuestion: {question}"
            )
            model = FakeListChatModel(responses=["(the model's grounded answer)"])
            
            
            def format_docs(found):
                return "\n\n".join(d.page_content for d in found)
            
            
            rag = (
                {"context": retriever | format_docs, "question": RunnablePassthrough()}
                | prompt
                | model
                | StrOutputParser()
            )
            print(rag.invoke("What does LangGraph add?"))

            The fake embeddings here are random vectors, so the ranking means nothing; swap in a real embedding model and the code stays the same.

            Chain, create_agent, or a custom graph?

            Pick the simplest shape that fits. Each shape builds on the one before it, and all three are Runnables, so you can move up later without rewriting your tools or prompts.

            1. Are the steps fixed, with no loop?
              A chain (LCEL)

              Prompt, model, parser, maybe a retriever. Predictable, cheap and easy to test. Most extraction, classification, summarising and RAG apps are chains.

              prompt | model | parser
            2. Does the model choose tools until it’s done?
              create_agent

              The standard tool-calling loop, built on LangGraph. Add behaviour with middleware instead of rewriting the loop: human approval, summarising long histories, retries, fallback models, call limits, PII redaction.

              create_agent(model, tools, middleware=[...])
            3. Do you need your own control flow?
              A custom StateGraph

              Several agents, branches, parallel fan-out, approval at a specific step, long-running jobs that must survive restarts. You design the state, the nodes and the edges.

              StateGraph(State).add_node(...)

            create_agent builds the agent loop for you

            The same fake model and tool as in the stepper. The result is a compiled LangGraph graph with nodes model and tools. The system_prompt is sent with every model call but is not stored in the state.

            from langchain.agents import create_agent
            from langchain_core.language_models import GenericFakeChatModel
            from langchain_core.messages import AIMessage, HumanMessage
            from langchain_core.tools import tool
            
            
            class FakeChatModel(GenericFakeChatModel):
                def bind_tools(self, tools, **kwargs):
                    return self
            
            
            @tool
            def get_weather(city: str) -> str:
                """Get the current weather for a city."""
                return f"Sunny, 21 C in {city}"
            
            
            call = {"name": "get_weather", "args": {"city": "Paris"}, "id": "call_1"}
            model = FakeChatModel(
                messages=iter([AIMessage("", tool_calls=[call]), AIMessage("Sunny, 21 C.")])
            )
            
            agent = create_agent(model, tools=[get_weather], system_prompt="You report weather.")
            result = agent.invoke({"messages": [HumanMessage("Weather in Paris?")]})
            print(list(agent.get_graph().nodes))
            print([m.type for m in result["messages"]])
            ['__start__', 'model', 'tools', '__end__']
            ['human', 'ai', 'tool', 'ai']

            Middleware: hooks around the loop

            Middleware wraps the model and tool calls inside create_agent. Built-in ones, all in langchain.agents.middleware:

            • HumanInTheLoopMiddleware: pause before chosen tools; a person can approve, edit, reject or respond.
            • SummarizationMiddleware: summarise old messages when the history gets long.
            • ModelRetryMiddleware, ToolRetryMiddleware, ModelFallbackMiddleware: retry with backoff, then switch models.
            • ModelCallLimitMiddleware, ToolCallLimitMiddleware: cap calls per run or per thread.
            • PIIMiddleware: redact or block personal data.

            Write your own with the decorators @before_model, @after_model, @wrap_model_call, @wrap_tool_call and @dynamic_prompt.

            And when you need neither

            One model call with no tools, no retrieval and no steps? The provider’s SDK, or a single init_chat_model(...).invoke(...), is enough. Frameworks pay off when you swap providers, compose steps, stream, trace, or run loops with state. Don’t add layers you can’t explain.

            LangGraph core concepts

            A LangGraph app is a state (a TypedDict, dataclass or Pydantic model), nodes that read it and return updates, and edges that decide which node runs next. It runs in super-steps: all nodes scheduled for a step run (in parallel if there are several), their updates are merged, then the next step starts.

            Streaming modes

            streams.py one node, one model call
              Output

                State and reducers

                Nodes return partial updates, never the whole state. A key without a reducer is overwritten; a key annotated with a reducer is merged. add_messages appends new messages and replaces one whose id matches, which is how you edit a message.

                import operator
                from typing import Annotated, TypedDict
                
                from langchain_core.messages import AIMessage, HumanMessage
                from langgraph.graph import add_messages
                
                
                class State(TypedDict):
                    question: str  # no reducer: a new value overwrites the old one
                    log: Annotated[list[str], operator.add]  # reducer: lists are concatenated
                    messages: Annotated[list, add_messages]  # append, or replace by message id
                
                
                print(operator.add(["start"], ["agent"]))
                
                old = [HumanMessage("Hi", id="1"), AIMessage("Draft", id="2")]
                print([m.content for m in add_messages(old, [AIMessage("Final", id="2")])])
                print([m.content for m in add_messages(old, [AIMessage("More")])])
                ['start', 'agent']
                ['Hi', 'Final']
                ['Hi', 'Draft', 'More']

                MessagesState is the ready-made state with just messages and add_messages; subclass it to add keys.

                Nodes, edges and conditional edges

                • A node is a function (sync or async) that takes the state and returns a dict of updates.
                • add_edge(a, b) always goes from a to b. START and END mark the entry and exit.
                • add_conditional_edges(a, router) calls router(state) and goes where it returns. Pass a list or dict of targets so the graph can be drawn.
                • A node can also return Command(goto="b", update={...}) to update state and route in one place.
                • An edge back to an earlier node makes a cycle. That loop is the difference between a graph and a chain.

                Cycles and the recursion limit

                Every super-step counts towards recursion_limit. When a run hits it, LangGraph raises GraphRecursionError instead of looping forever. Set it per call in the config.

                from typing import TypedDict
                
                from langgraph.errors import GraphRecursionError
                from langgraph.graph import START, StateGraph
                
                
                class State(TypedDict):
                    n: int
                
                
                builder = StateGraph(State)
                builder.add_node("loop", lambda s: {"n": s["n"] + 1})
                builder.add_edge(START, "loop")
                builder.add_edge("loop", "loop")  # a cycle with no exit
                graph = builder.compile()
                
                try:
                    graph.invoke({"n": 0}, {"recursion_limit": 5})
                except GraphRecursionError as e:
                    print(type(e).__name__, str(e).splitlines()[0])
                GraphRecursionError Recursion limit of 5 reached without hitting a stop condition. You can increase the limit by setting the `recursion_limit` config key.

                Older releases defaulted to 25 steps. In langgraph 1.2 the default is 10,007 (LANGGRAPH_DEFAULT_RECURSION_LIMIT), and create_agent sets 9,999, so set your own limit for agents that could loop.

                Checkpointers, threads and time travel

                Compile with a checkpointer and every super-step is saved as a checkpoint under the thread_id in the config. That gives you conversation memory, resumable runs after a crash, interrupts, and time travel: go back to an old checkpoint, change it, and run forward on a new branch.

                from typing import TypedDict
                
                from langgraph.checkpoint.memory import InMemorySaver
                from langgraph.graph import START, StateGraph
                
                
                class State(TypedDict):
                    n: int
                
                
                builder = StateGraph(State)
                builder.add_node("double", lambda s: {"n": s["n"] * 2})
                builder.add_node("inc", lambda s: {"n": s["n"] + 1})
                builder.add_edge(START, "double")
                builder.add_edge("double", "inc")
                graph = builder.compile(checkpointer=InMemorySaver())
                
                config = {"configurable": {"thread_id": "t"}}
                print(graph.invoke({"n": 5}, config))
                
                history = list(graph.get_state_history(config))  # newest first
                before_inc = next(c for c in history if c.next == ("inc",))
                print(before_inc.values)
                
                fork = graph.update_state(before_inc.config, {"n": 100})
                print(graph.invoke(None, fork))
                {'n': 11}
                {'n': 10}
                {'n': 101}

                Use InMemorySaver in tests and a database saver in production, such as PostgresSaver from langgraph-checkpoint-postgres.

                Interrupts

                • interrupt(value) inside a node pauses the run and returns value to the caller under __interrupt__. It needs a checkpointer.
                • Resume with graph.invoke(Command(resume=answer), config) on the same thread. The node restarts from its first line, and interrupt() returns answer.
                • Because of the restart, keep side effects after the interrupt() call, or make them idempotent.
                • interrupt_before and interrupt_after at compile time pause around whole nodes; they are mostly for debugging.

                Send: map-reduce and parallel branches

                When the number of branches is only known at run time, return a list of Send(node, input) from a conditional edge. Each one runs the node with its own input, in parallel in the same super-step, and a reducer collects the results.

                import operator
                from typing import Annotated, TypedDict
                
                from langgraph.graph import START, StateGraph
                from langgraph.types import Send
                
                
                class State(TypedDict):
                    topics: list[str]
                    summaries: Annotated[list[str], operator.add]
                
                
                def fan_out(state: State):
                    return [Send("summarise", {"topic": t}) for t in state["topics"]]
                
                
                def summarise(item: dict):
                    return {"summaries": [f"summary of {item['topic']}"]}
                
                
                builder = StateGraph(State)
                builder.add_node("summarise", summarise)
                builder.add_conditional_edges(START, fan_out, ["summarise"])
                graph = builder.compile()
                
                print(graph.invoke({"topics": ["cats", "dogs", "owls"]}))
                {'topics': ['cats', 'dogs', 'owls'], 'summaries': ['summary of cats', 'summary of dogs', 'summary of owls']}

                For a fixed set of branches you don’t need Send: add several edges out of one node and the targets run in parallel.

                Subgraphs

                A compiled graph is a Runnable, so it can be a node in another graph. If both share state keys, pass it straight to add_node. If the schemas differ, call it inside a node function and map the state in and out. Subgraphs are how you build multi-agent systems from smaller, testable agents.

                from typing import TypedDict
                
                from langgraph.graph import START, StateGraph
                
                
                class State(TypedDict):
                    text: str
                
                
                def clean(state: State):
                    return {"text": state["text"].strip()}
                
                
                inner = StateGraph(State)
                inner.add_node("clean", clean)
                inner.add_edge(START, "clean")
                cleaner = inner.compile()
                
                outer = StateGraph(State)
                outer.add_node("cleaner", cleaner)  # a compiled graph is a node
                outer.add_node("shout", lambda s: {"text": s["text"].upper()})
                outer.add_edge(START, "cleaner")
                outer.add_edge("cleaner", "shout")
                print(outer.compile().invoke({"text": "  hi  "}))
                {'text': 'HI'}

                Long-term memory: the Store

                A checkpointer remembers one thread. A Store keeps JSON documents under namespaces, shared across threads: user preferences, learned facts. Nodes reach it, and the per-run context, through the Runtime argument.

                from dataclasses import dataclass
                
                from langgraph.graph import START, MessagesState, StateGraph
                from langgraph.runtime import Runtime
                from langgraph.store.memory import InMemoryStore
                
                
                @dataclass
                class Context:
                    user_id: str
                
                
                def remember(state: MessagesState, runtime: Runtime[Context]):
                    ns = ("users", runtime.context.user_id)
                    runtime.store.put(ns, "prefs", {"language": "Romanian"})
                    return {}
                
                
                builder = StateGraph(MessagesState, context_schema=Context)
                builder.add_node("remember", remember)
                builder.add_edge(START, "remember")
                store = InMemoryStore()
                graph = builder.compile(store=store)
                
                graph.invoke({"messages": []}, context=Context(user_id="ana"))
                print(store.get(("users", "ana"), "prefs").value)
                {'language': 'Romanian'}

                With an embedding index configured, store.search(namespace, query=...) finds memories by meaning.

                Production essentials

                The parts interviewers ask about once your demo works: seeing what happened, surviving flaky providers, stopping runaway loops, and testing without paying for tokens.

                LangSmith tracing

                Set LANGSMITH_TRACING=true and LANGSMITH_API_KEY (and optionally LANGSMITH_PROJECT) and every chain, agent and graph run is recorded as a tree: each step’s inputs, outputs, latency, token usage and errors. No code changes. Wrap your own functions with @traceable from the langsmith package to include them.

                LangSmith evaluation

                Keep a dataset of example inputs and reference outputs, run your app over it with evaluate(target, data=..., evaluators=[...]), and score each result with code checks or an LLM judge. Each run is an experiment, so you can compare prompts and models side by side before you ship, then run online evaluators on production traces.

                Retries, fallbacks and rate limits

                with_retry retries a Runnable with exponential backoff; with_fallbacks tries the next Runnable when it still fails. Here the primary always times out:

                from langchain_core.language_models import FakeListChatModel
                from langchain_core.runnables import RunnableLambda
                
                attempts = []
                
                
                def flaky_model(prompt):
                    attempts.append(prompt)
                    raise TimeoutError("provider timed out")
                
                
                primary = RunnableLambda(flaky_model)
                backup = FakeListChatModel(responses=["Answer from the backup model"])
                
                model = primary.with_retry(stop_after_attempt=3).with_fallbacks([backup])
                print(model.invoke("hi").content)
                print(len(attempts))
                Answer from the backup model
                3

                To stay under a provider’s rate limit, pass rate_limiter=InMemoryRateLimiter(requests_per_second=...) to the chat model. In agents, use the retry and fallback middleware.

                Limits and long conversations

                • Set recursion_limit, and in agents ModelCallLimitMiddleware or ToolCallLimitMiddleware, so a confused model can’t loop and spend forever.
                • Histories outgrow the context window: trim them with trim_messages, or summarise them with SummarizationMiddleware.
                • Use a database checkpointer and choose durability ("sync", "async" or "exit") to trade safety for speed.

                Testing with fake models

                FakeListChatModel replays strings, GenericFakeChatModel and FakeMessagesListChatModel replay whole AIMessages, including tool_calls. Script the model, run the real chain or graph, and assert on the messages, the route taken and the final state, exactly like every program on this page. Unit tests stay fast, free and deterministic; check answer quality separately with LangSmith evaluations.

                LangChain, LangGraph, LangSmith

                • LangChain: the parts (models, messages, tools, retrievers), LCEL, and create_agent.
                • LangGraph: the runtime for stateful, looping, durable workflows. create_agent runs on it.
                • LangSmith: tracing, evaluation and monitoring. It works with or without the other two.

                What to remember

                Everything is a Runnableinvoke, batch, stream and their async twins work on every part. | builds a RunnableSequence; a dict becomes a RunnableParallel.
                Messages are the interfaceSystem, Human, AI and Tool messages. An AIMessage can ask for tools; each ToolMessage answers one call by tool_call_id.
                The model never runs toolsIt returns tool_calls. Your code, ToolNode or create_agent runs them and sends the results back.
                An agent is a loopModel, tools, model, until there are no tool calls. create_agent builds it; StateGraph lets you shape it.
                State plus reducersNodes return partial updates and reducers merge them. add_messages appends, or replaces by id.
                Checkpointer plus thread_idThat pair gives you memory, interrupts with Command(resume=...), crash recovery and time travel. A Store remembers across threads.

                Interview questions: LangChain and LangGraph

                Short answers you can say out loud. Try answering each question yourself before you open it.

                What is LangChain for, and when would you not use it?

                LangChain gives you standard, provider-independent parts for LLM apps: chat models, messages, prompts, tools, retrievers and output parsing, all composable as Runnables, plus create_agent for tool-calling agents. It pays off when you compose steps, swap providers, stream, trace or build agents. For a single model call with no tools or retrieval, the provider’s SDK is simpler, and you shouldn’t add a layer you can’t explain.

                What is the difference between LangChain, LangGraph and LangSmith?

                LangChain is the set of building blocks and the high-level create_agent. LangGraph is the low-level runtime for stateful workflows: graphs with state, cycles, checkpoints, interrupts and streaming; create_agent runs on it. LangSmith is the observability and evaluation platform, and it works with or without the other two.

                What is a Runnable, and what is LCEL?

                A Runnable is anything with the standard interface: invoke, batch, stream and their async versions. Prompts, models, parsers, retrievers, tools and compiled graphs are all Runnables. LCEL, the LangChain Expression Language, composes them with |, which builds a RunnableSequence where each step’s output is the next step’s input. The composed chain is itself a Runnable, so it gets streaming, batching, async and tracing for free.

                What is the difference between invoke, batch, stream and astream?

                invoke takes one input and returns one output. batch takes a list and runs the inputs concurrently on a thread pool, limited by max_concurrency. stream yields the output in chunks, such as tokens, as soon as they are ready. astream and the other a methods are the async versions for asyncio servers, so a slow model call doesn’t block other requests.

                What do RunnableParallel, RunnablePassthrough and RunnableLambda do?

                RunnableParallel runs several Runnables on the same input and returns a dict of their results; a plain dict inside a chain is turned into one automatically. RunnablePassthrough passes its input through unchanged, and RunnablePassthrough.assign adds computed keys to an input dict. RunnableLambda wraps a plain Python function so it can sit in a chain. The classic RAG chain uses all three ideas: {"context": retriever | format, "question": RunnablePassthrough()} | prompt | model.

                What are the message types, and what does each one mean?

                SystemMessage carries instructions, HumanMessage the user’s input, and AIMessage the model’s reply, which can include tool_calls and usage_metadata. ToolMessage carries a tool’s result and must have the tool_call_id of the call it answers. A chat model takes a list of messages and returns an AIMessage; in 1.x, .text gives the text and .content_blocks gives typed blocks in the same shape for every provider.

                How does tool calling work, step by step?

                You bind tools with model.bind_tools(tools), which sends each tool’s name, description and JSON schema with the request. The model doesn’t run anything: it returns an AIMessage with tool_calls, each with a name, args and an id. Your code runs each tool and appends a ToolMessage with the matching tool_call_id, then calls the model again with the whole history. When the model replies without tool calls, you’re done.

                bind_tools or with_structured_output: when do you use each?

                Use bind_tools when the model should decide whether to act and which tools to call, as in an agent. Use with_structured_output(Schema) when you always want data in a fixed shape, such as extraction or classification: it forces the format with the provider’s JSON mode or a tool call, then parses the result into your Pydantic model, TypedDict or dict. include_raw=True also returns the raw message and any parsing error.

                How does an agent loop work?

                The model gets the conversation and the tool schemas. If it returns tool calls, the tools run and their results are appended as ToolMessages, then the model is called again. The loop ends when the model answers without tool calls, or when a limit stops it. In LangGraph that’s two nodes, model and tools, a conditional edge such as tools_condition, and an edge from tools back to the model.

                When would you use create_agent, and when a hand-built LangGraph graph?

                create_agent(model, tools, ...) gives you the standard tool-calling loop on LangGraph, with checkpointing, streaming and middleware for things like human approval, summarisation, retries and call limits. Use it whenever the model choosing tools in a loop is the right shape. Build your own StateGraph when you need custom control flow: fixed steps mixed with agent steps, several agents, parallel fan-out, or approval at one specific step.

                What is middleware in create_agent?

                Hooks that run around the agent’s model and tool calls, so you change behaviour without rewriting the loop. Built-in ones include HumanInTheLoopMiddleware, SummarizationMiddleware, ModelRetryMiddleware, ModelFallbackMiddleware, ModelCallLimitMiddleware, ToolCallLimitMiddleware and PIIMiddleware. You write your own with decorators such as @before_model, @after_model, @wrap_model_call and @dynamic_prompt.

                What are StateGraph, state and reducers? Why add_messages?

                A StateGraph is built from a state schema, usually a TypedDict. Nodes return partial updates, and each key’s reducer decides how an update is merged: no reducer means overwrite, operator.add means concatenate. add_messages appends new messages and replaces any with the same id, so nodes can return just the new message and parallel nodes don’t overwrite each other. MessagesState is the ready-made state with a messages key using it.

                How do conditional edges and cycles work?

                add_conditional_edges(node, router) calls router(state) after the node and goes to the node it returns, or to END. An edge back to an earlier node creates a cycle, which is how agents loop. A node can also return Command(goto=..., update=...) to route and update in one step. Every loop needs an exit condition, and the recursion limit is the safety net.

                What is the recursion limit, and what happens when you hit it?

                It caps the number of super-steps in one run. When it’s reached, LangGraph raises GraphRecursionError instead of looping forever. Set it per call with config={"recursion_limit": n}. Older versions defaulted to 25; langgraph 1.2 defaults to 10,007 and create_agent uses 9,999, so set your own for agents, or add call-limit middleware.

                What does a checkpointer do, and what is a thread?

                A checkpointer saves the graph’s state after every super-step. A thread is one conversation or job, identified by thread_id in config["configurable"]; all its checkpoints are stored under that id. Invoke again with the same thread and the graph continues from the saved state, which gives you conversation memory, recovery after a crash, interrupts and time travel. Use InMemorySaver in tests and a Postgres or SQLite saver in production.

                What is time travel in LangGraph?

                Because every step is checkpointed, you can list past states with get_state_history(config). Invoking with an older checkpoint’s config replays from that point. update_state on an old checkpoint creates a fork you can run forward with different values. It’s used for debugging, for retrying from a bad step, and for letting users edit an earlier turn.

                How do interrupt and Command(resume=...) work, and what’s the catch?

                Calling interrupt(value) in a node saves the state and stops the run; the caller gets value under __interrupt__. To continue, invoke the same thread with Command(resume=answer). The catch: the node restarts from its first line and interrupt() then returns answer, so any code before it runs twice. Keep side effects after the interrupt or make them idempotent. Interrupts need a checkpointer.

                Short-term or long-term memory: checkpointer or Store?

                Short-term memory is the state of one thread, kept by the checkpointer: the messages of this conversation. Long-term memory has to outlive threads, such as user preferences or facts learned earlier, so it goes in a Store: JSON documents under namespaces like ("users", user_id), optionally searchable by meaning. Nodes reach the store through the Runtime argument.

                What is Send for, and how do parallel branches work?

                Nodes scheduled in the same super-step run in parallel, so several edges out of one node already give you fixed parallel branches. When the number of branches is only known at run time, a conditional edge returns a list of Send(node, input), and each one runs the node with its own input, the map step of map-reduce. A reducer such as operator.add on the result key collects their outputs.

                What are subgraphs, and why use them?

                A compiled graph is a Runnable, so it can be a node inside another graph. If the parent and child share state keys, you add it directly; if not, you call it inside a node and map the state in and out. Subgraphs let you build and test agents separately and then combine them, which is the usual way to build multi-agent systems. A node in a subgraph can route in the parent with Command(graph=Command.PARENT, goto=...).

                What streaming modes does LangGraph have?

                values streams the full state after each step, and updates streams only what each node returned. messages streams LLM tokens with metadata saying which node produced them, and custom streams whatever nodes send with get_stream_writer(). There are also checkpoints, tasks and debug; pass a list to get several as (mode, data) tuples.

                How do you stream an agent’s answer to a web UI?

                Use astream with stream_mode="messages" to get tokens as they’re generated, often combined with "updates" to show progress such as which tool is running. Forward the chunks to the browser over server-sent events or a WebSocket. Token streaming works even if the node calls invoke, because LangGraph listens to the model’s callbacks.

                How do you build RAG with LangChain?

                Indexing: load sources into Documents, split them into chunks with a text splitter, embed the chunks and store them in a vector store. Answering: a retriever finds the most similar chunks for the question, and a chain puts them into the prompt and asks the model to answer from that context only. For harder questions you make retrieval a tool, so an agent can search several times and rewrite its query.

                How do you choose chunk size and overlap?

                Chunks should be big enough to hold one complete idea and small enough that a retrieved chunk is mostly relevant; a few hundred to a thousand or so tokens is a common start. Overlap, often 10 to 20 percent, repeats text across neighbouring chunks so an idea cut at a boundary is still found. RecursiveCharacterTextSplitter splits on paragraphs, then lines, then words, to keep chunks natural. Then measure retrieval quality on real questions and tune.

                What is a retriever, and how is it different from a vector store?

                A vector store stores embeddings and runs similarity search. A retriever is a Runnable interface: a query string in, a list of Documents out. vector_store.as_retriever(search_kwargs={"k": 4}) wraps a store, but a retriever can also be keyword search, a web search or a hybrid of several. Because it’s a Runnable, it drops into any chain.

                What is LangSmith used for?

                Tracing: with LANGSMITH_TRACING=true and an API key, every run is recorded as a tree of steps with inputs, outputs, latency, tokens and errors, so you can see exactly what the model was sent. Evaluation: you keep datasets of examples, run your app over them with code or LLM-judge evaluators, and compare experiments before shipping. It also monitors production traffic and can run online evaluations on it.

                How do you handle retries, fallbacks and rate limits?

                runnable.with_retry(stop_after_attempt=3) retries with exponential backoff, and with_fallbacks([backup]) tries another model or chain when it still fails. Pass rate_limiter=InMemoryRateLimiter(requests_per_second=...) to a chat model to stay under provider limits. In agents, ModelRetryMiddleware, ToolRetryMiddleware and ModelFallbackMiddleware do the same jobs.

                How do you test LangChain and LangGraph code without calling a real model?

                Use the fake chat models in langchain_core: FakeListChatModel replays strings, and GenericFakeChatModel or FakeMessagesListChatModel replay whole AIMessages, including tool calls. Run the real chain or graph with an InMemorySaver and assert on the messages, the route taken and the final state. That keeps unit tests fast and deterministic; you judge answer quality separately with LangSmith evaluations.

                How do you handle a conversation that no longer fits in the context window?

                Trim it, keeping the system message and the most recent messages, with trim_messages; summarise older turns into a short message, which SummarizationMiddleware does for agents; or move lasting facts to a Store and retrieve them when needed. When trimming, never split an AIMessage with tool calls from its ToolMessages, or the provider rejects the history.

                What happens when a tool raises an exception?

                In a graph, ToolNode by default turns invalid-argument errors into a ToolMessage so the model can correct its call, and re-raises other exceptions. Set handle_tool_errors=True, or pass a function, to send every error back to the model as a message. In create_agent, retry middleware can retry flaky tools first.

                What does init_chat_model do?

                It creates a chat model from a string like "provider:model-name", loading the right integration package, such as langchain-openai or langchain-anthropic. Every provider then has the same interface, so you can switch models through configuration without changing code. create_agent also accepts the same string directly.

                How do multi-agent systems work in LangGraph?

                Each agent is a node or a subgraph with its own prompt and tools. In a supervisor design, one agent routes work to the others, often by calling them as tools; in a handoff design, agents pass control directly with Command(goto=...). Shared state or messages carry the context between them. Start with one agent and add more only when tools or instructions clearly separate.

                Every output on this page was printed by the programs in verify/langchain/, run with langchain 1.4.3, langchain-core 1.6.7 and langgraph 1.2.14 on Python 3.14. The models are the fake chat models from langchain_core.language_models; message ids are left out of the drawings.