Tools and resources

How I work with AI.

I try a lot of AI software. This is the smaller set that has stayed in my workflow. I delegate plenty, but I keep the checks visible and make the software enforce the boundaries that matter.

More than one agent, on purpose.

Most changes begin as a conversation with an agent and end with me reviewing the diff. If the problem matters, I will often hand it to a second agent before I trust the first answer.

Claude Code

Claude Code is where most of my building starts. I use it from the desktop app so I can keep several local or remote sessions moving at once. It can read the repository, edit files, run commands, and test the change. When the interface needs work, I can bring Claude Design into the same loop.

anthropic.com/claude-code ↗

Cowork

Cowork gets the work that belongs in folders instead of repositories. I use it for spreadsheets, documents, training material, and jobs that pull information from several sources. It works directly in the folders I choose, so the result comes back as a file I can review instead of a long chat I have to reassemble.

anthropic.com/product/claude-cowork ↗

Codex

Codex is my second coding agent, and I run it from the terminal. I often hand the same problem to Codex and Claude Code, then compare their approaches. Agreement is useful, but a disagreement is usually where I find the assumption that needs another look.

developers.openai.com/codex ↗

ChatGPT Work

I use ChatGPT Work for deliverables such as documents, decks, spreadsheets, and small sites. It is also a second read on work I might give Cowork. The workspace can work from source files, produce an editable artifact, and call Codex when the deliverable needs code.

learn.chatgpt.com ↗

Skills, tools, and plugins I use.

Below are some of my favorite skills, plugins, and MCP servers that I use to make building applications more productive.

Stripe MCP

I use Stripe's MCP server while building payment and billing flows. It can search Stripe's documentation and work with Stripe API data, so I can inspect an account and make approved changes without bouncing between the editor, API reference, and Dashboard.

docs.stripe.com/mcp ↗

Context7 MCP

I use Context7 to pull current, version-specific documentation and code examples into the agent's context. It resolves the library name and finds the relevant docs before the agent starts writing code. That helps when a framework or package may have changed since the model was trained.

context7.com ↗

Playwright CLI

Playwright CLI gives coding agents such as Claude Code and Codex direct browser control from the command line. They can navigate and interact with a site, inspect page snapshots, take screenshots, and record traces or video. I use it for automated testing and automated documentation, where the same browser run can verify a workflow and capture the screens needed to explain it.

playwright.dev/agent-cli/introduction ↗

Engineering skills

Matt Pocock's engineering skills cover the work around the code: shaping a plan, test-driven development, domain language, and code review. I use Grilling most. It asks one question at a time and keeps pressing until a fuzzy plan becomes explicit. The process is mildly irritating in exactly the right way, and it catches decisions I was about to skip. Matt is also a brilliant teacher. His popular YouTube videos on agentic AI development are worth watching if you want to see how these ideas work in practice.

github.com/mattpocock/skills ↗

Ponytail

An agent can turn a small problem into an architecture. Ponytail is the brake pedal I use when it starts doing that. It pushes YAGNI, the standard library before another dependency, and one clear function before a new abstraction. The result is usually less code for me to review and maintain.

github.com/DietrichGebert/ponytail ↗

Impeccable

I use Impeccable after an interface exists. It audits the frontend for accessibility problems, weak design choices, and the habits that make AI-generated UI look interchangeable. Its setup step also records what the product is and who it is for. That gives later design work a factual place to start.

impeccable.style ↗

Claude Design

I use Claude Design when a page needs visual work, not another paragraph describing visual work. It can import a codebase or existing design system, build an interactive prototype, and hand the result back to Claude Code. I would rather react to something I can click than debate a static mockup.

claude.com/product/design ↗

Graphify

On a large or unfamiliar repository, I use Graphify to give the agent a narrower place to start. It builds a knowledge graph from local AST parsing, with no embeddings or vector store, and returns a traceable subgraph for a question. That cuts down on blind searching before the real work begins.

github.com/Graphify-Labs/graphify ↗

How I build agentic workflows.

Pydantic and Pydantic AI give my agents typed data, tools, and outputs. LangGraph is my preferred tool for coordinating those parts when the application needs a workflow I can inspect and control.

Pydantic

When I build an AI-enabled application, Pydantic gives my code a structured way to interact with the LLM. Models are useful, but their answers are not trustworthy data. They can leave out a field or return a value in a shape the application does not expect. Pydantic checks the response before anything else depends on it, so the application gets typed data or a validation error it knows how to handle.

pydantic.dev ↗

Pydantic AI

Pydantic AI is my preference for a focused Python AI application framework with typed dependencies, tools, and output. If the model returns the wrong shape, the framework can send the validation error back and ask it to try again. That gives the rest of the application a real contract to work against.

ai.pydantic.dev ↗

LangGraph

LangGraph is my preferred tool for agent orchestration. It lets me model a workflow as explicit state and transitions, then persist that state when the work has to pause or recover from a failure. I reach for it over Pydantic AI when I need to coordinate several agents or tools and put human approval into the flow. It also lets me mix predictable application code with decisions made by a model. Each step remains visible in the graph.

langchain.com/langgraph ↗

What carries over

What I expect
from the stack.

When I build with AI, I need a clear trail back to the data the model used. That lets me check that the result is grounded in its source material before I trust it. The rest of my working stack includes Next.js, React Native and Expo, Vercel, Supabase, Neon and Playwright, Azure OpenAI, and sometimes Gemini for its vision capabilities.