Hype StackHypeStack

Search

Search the packs, templates, docs, and pages

Guide

Cursor Best Practices From Running an Agent on a Large Codebase

Most Cursor advice is about the prompt. On a large repository the results depend more on what the agent can check, what it is told not to run, and what it can look up instead of inventing.

The usual list of cursor best practices is short and mostly right: use Plan Mode, keep the context small, review the diff, commit often. If you have been using Cursor on a real project for a few weeks you already do those, and the results are still uneven. One session lands a clean feature. The next renames a concept, passes its own tests, and breaks the type check in a package it did not open.

The habits below come from running an agent every day on a monorepo with several apps and shared packages. Each one exists because something went wrong without it. They are written for Cursor's mechanics: rules, Plan Mode, checkpoints, run modes, hooks and cloud agents.

If you use Cursor mainly for tab completion and single-file edits, skip most of this. It pays back when the agent makes multi-file changes you do not read line by line.

Make the agent prove the change, with more than tests

The most expensive failure we have seen is a green test run that meant nothing. After a test-runner upgrade, tests passed at runtime while the type check failed on every assertion, because the matcher typings had drifted. The agent ran the tests, saw green, and reported the work done.

So the finish line is three commands, in a rule the agent sees in every session: lint the files it edited, run the type check for the app it touched, then run the tests. In Cursor that is a project rule set to Always Apply, kept to a few lines because it is paid for in every chat. Name the exact commands.

What the finish line checks

Lint, on the edited files

Seconds to run, and it catches the unused import and the floating promise before a reviewer has to. Scope it to the files the agent touched so it runs every turn.

Type check, for the whole app

It sees files the agent did not open. A changed response shape fails here in the three consumers nobody mentioned in the prompt.

Tests, last

Tests prove the behaviour the agent thought about. Run them after the type check, because a red type check often explains the test failure for free.

For dependency and tooling bumps, run all three from the repository root. An app-scoped check does not see the packages that share the tooling.

Cursor's hooks can enforce the same thing without relying on the model's memory. A project-level .cursor/hooks.json can run a script after a file edit or when the agent stops, which is a reasonable place for a linter.

Get a failing repro before any fix

Ask Cursor to fix a bug and it will read the code, form a theory, and edit three files. You cannot tell from the diff whether the theory was right.

The rule we use: before touching the fix, the agent writes a test that fails on this bug, and runs it red once. Then one hypothesis at a time against that test. The fix is done when the test goes green and the surrounding suite still passes, and the test stays in the codebase as the regression test.

This is where Cursor's checkpoints mislead people. A checkpoint snapshots files during an agent session and lets you undo the agent's edits. It is stored locally, separate from Git, and restoring one reverts files without removing the conversation. That is an undo button, not a record of what was tried. Commit the failing test on its own before the agent starts on the fix. If the session goes sideways, you restore to a commit that still contains the repro.

Give it a glossary and a decision log

Two failures look like carelessness and are really missing information.

The first is synonyms. Over a month, an agent with no vocabulary to follow will call the same thing an account, a user, and a member, in three tables and two route paths. A one-page glossary at the repository root fixes it: one line per term, with the definition. Cursor's own guidance for rules is to reference files instead of copying their contents, so the always-on rule is two sentences: use the terms in the glossary exactly, and add a line when you introduce a concept.

The second is the helpful undo. You chose a query builder over the ORM for one hot path, and three weeks later the agent tidies it back. Short decision records stop that: context, decision, consequences, a few paragraphs each, with a rule telling the agent to read the records for the area before proposing structural changes and to ask before working around one.

Plan Mode fits here. It researches the codebase, asks clarifying questions and writes a plan you can edit before anything is built. Plans are saved outside the repository by default, and there is a "Save to workspace" option. Use it for the plans that contain a decision, so the reasoning sits in the repository for the next session to read.

Say what not to run, and gate what leaves

Agents start things. A dev server already running in your terminal gets started again on another port, and the session spends minutes on a build that only existed to see whether the page renders.

Four boundaries worth writing down

  • A banned-commands rule, with the alternativeList the dev, serve and container commands the agent must leave alone, state which ports the running apps answer on, and say that tests are fine to run. A prohibition without a replacement gets worked around.
  • A run mode you chose on purposeCursor's run modes decide which commands run without asking, which go to the sandbox, and which prompt you. Put lint, type check and test commands on the allowlist so verification costs no clicks, and leave the rest reviewed.
  • A hook for the commands a rule cannot be trusted withA rule is a request. A beforeShellExecution hook can deny a command outright, which is the right tool for a push to the main branch or a destructive database command.
  • A push gate that matches the changeA copy edit needs that app's lint. A dependency bump needs the root type check, the full test run and an end-to-end pass. Write the table down so the agent runs the checks the change can break and skips the rest.

Cloud agents make the gate more important. They clone the repository into their own machine, work on a separate branch and push it for handoff, so the branch arrives without anyone having watched it being written.

Cloud agents also change what "do not start the server" means. Locally the servers are yours. In a cloud agent's machine nothing is running, and the agent is only as capable as its environment, which Cursor lets you define in .cursor/environment.json. Scope the banned-commands rule to local work, or the cloud agent will refuse to boot the one thing it needs.

A rule asks the agent to behave. A type check, a failing test and a denied command do not depend on it remembering.

Where a foundation helps, and where it does not

Every practice above assumes the repository has something to check against: a type check that crosses the API boundary, a test setup the agent can run, conventions written down. On a new project you can start with those in place instead of writing them in month three.

That is the part we make. Hype Stack is a full-stack TypeScript template, and its agent rules are authored as Cursor .mdc files in .cursor/rules, then converted to other editors' formats when you ask for them. Cursor is the default editor in create, so a new project gets the rules as written: an always-on set, and scoped ones that load only when the agent opens a matching path. They cover backend error handling, feature structure, testing workflow and verification after changes. create also installs a starter set of skills into .agents/skills. The MCP server lets Cursor add accounts, billing or a team workspace through the real installer, and npx @hype-stack/cli@latest mcp install --editor cursor writes the entry into ~/.cursor/mcp.json.

What is not in the box matters more. Rules files do not make an agent correct. They are conventions for one stack. They do nothing for a codebase that is not built on it, and copying them into an existing repository gives the agent confident instructions about files you do not have. The template does not ship your glossary, your decision records or your push gate either, because those describe your product. On an existing codebase, the practices on this page are the work, and no template shortens it.

Templates with the Cursor rules already in .cursor/rulesSix working apps, each with path-scoped rules, a typed contract from backend to frontend, and an MCP server Cursor can call to add packs. The glossary and the decision log are yours to write.Browse templates

What you write yourself, on any foundation: the glossary, the decisions, the banned commands, and the table of checks each kind of change needs. A starting point with the rules in place follows below.

Hype Stack

What is Hype Stack?

Every product starts with the same month of work nobody pays you for: sign-up and login, teams and permissions, taking payments, notifications, an admin panel to run the business. Hype Stack is that month, already built and tested. Start from the free open-source app, add the pieces you need with one command, and keep going on the part that is actually your idea.

Everything lands as real code in your own repository, so there is nothing to rent and nothing anyone can switch off. For the engineers: React 19, Hono, Postgres, and a desktop build, typed end to end.

See it running

Templates are curated project starters built on this stack: a layout, feature packs, a custom theme, and bonus pages. The previews below are recordings of the real apps.

Two ways to install features

Same features underneath, different starting point. Either command resolves what the packs depend on, copies the source into your repository, and merges the Prisma schema.

Ready-made

Take a template

A landing page, a design system, a themed layout, and the features already wired into it. Rebrand it, put your product in the middle, ship.

$npx @hype-stack/cli template
Browse templates
From scratch

$ hype-stack compose

✓ Auth✓ Payments

Compose your own design

Your design and your choices, without rebuilding auth, billing, or notifications. Tick the packs you want and the CLI wires them into the open-source starter.

$npx @hype-stack/cli compose
Browse packs

Questions, answered

Make the agent finish with lint, a type check and tests, in that order, and name the exact commands in an always-on rule. The type check is the only one of the three that sees files the agent did not open, so it catches the consumers a changed response shape broke.

More stacks

Turn your ideas into
Real applications.

Start free and own every line you ship. When you want more, one All-Access license unlocks every premium pack and template for a year.

All premium packsEvery template12 months of updates
Get the whole catalog$299/year