โ† Browse

@itamarzand88/awesome-agent-conventions-12

A

A curated field guide to the convention files AI agents read, write, and act on.

instructionscodex

Install

agr install @itamarzand88/awesome-agent-conventions-12 --target codex

Writes 1 file into AGENTS.md, pinned to git-a022a858.

  • AGENTS.md

Document

AGENTS.md

This file provides guidance to AI agents when working with code in this repository.

๐Ÿšจ Legacy Codebase Guidelines ๐Ÿšจ

CRITICAL: This is a mature, mission-critical codebase (10+ years old in some areas). Follow these rules strictly:

Regression Prevention

  • NEVER introduce breaking changes without explicit approval
  • NEVER modify existing public APIs unless specifically requested
  • NEVER change existing test assertions or remove tests without approval
  • NEVER modify CI/build configuration without explicit request
  • ALWAYS run the complete test suite before committing changes
  • ALWAYS test against all supported Spark versions

Code Change Philosophy

  • Prefer addition over modification: Add new functionality alongside existing code
  • Preserve existing behavior: If unsure, err on the side of maintaining status quo
  • No "cowboy-style" refactoring: Avoid large-scale structural changes
  • Conservative by default: Only make changes that directly address the specific request

When in Doubt

  • Ask before changing: If a change seems risky, request explicit approval
  • Start small: Make minimal changes to achieve the goal
  • Maintain backward compatibility: Existing user code must continue to work
  • Document reasoning: Explain why changes are necessary in commit messages

Build and Development Commands

Core Build System

  • Build tool: sbt (Scala Build Tool) version 1.11.0
  • Build project: build/sbt compile
  • Run all tests: build/sbt test
  • Run single test suite: build/sbt "testOnly *PregelSuite"
  • Run single test: build/sbt "testOnly *PregelSuite -- -z 'test name pattern'"
  • Check code formatting: build/sbt scalafmtCheckAll
  • Apply code formatting: build/sbt scalafmtAll
  • Check code style: build/sbt "scalafixAll --check"
  • Generate documentation: build/sbt doc
  • Run with coverage: build/sbt "coverage test coverageReport"

Multi-Spark Version Testing

The project supports multiple Spark versions. Use -Dspark.version=X.Y.Z to specify:

  • build/sbt -Dspark.version=3.5.7 test
  • build/sbt -Dspark.version=4.0.1 test
  • build/sbt -Dspark.version=4.1.0 test

Pre-PR Checklist Commands

โš ๏ธ MANDATORY: These commands MUST pass before any PR. Failures indicate potential regressions.

Before raising a pull request, run these commands to ensure code quality and compatibility:

  1. Format code: build/sbt scalafmtAll
  2. Check formatting: build/sbt scalafmtCheckAll
  3. Check style/linting: build/sbt "scalafixAll --check"
  4. Build documentation: build/sbt doc
  5. Run tests with coverage: build/sbt "coverage test coverageReport"
  6. Test against multiple Spark versions:
    • build/sbt -Dspark.version=3.5.7 test
    • build/sbt -Dspark.version=4.0.1 test
    • build/sbt -Dspark.version=4.1.0 test

Alternative using pre-commit: pre-commit run all-files (handles Python formatting: black, isort, flake8)

Python Components

  • Python package location: /python/ directory
  • Python tests: Located in /python/tests/

Architecture Overview

Core Structure

GraphFrames is a graph processing library built on Apache Spark DataFrames with three main components:

  1. Core GraphFrame API (core/src/main/scala/org/graphframes/)

    • GraphFrame.scala: Main graph abstraction using DataFrames for vertices/edges
    • Provides graph algorithms, motif finding, and subgraph operations
  2. Algorithm Library (core/src/main/scala/org/graphframes/lib/)

    • Standard graph algorithms: PageRank, ConnectedComponents, ShortestPaths, etc.
    • Pregel.scala: Implements Pregel-style bulk synchronous parallel processing
    • AggregateMessages.scala: Lower-level message-passing API
  3. GraphX Compatibility Layer (graphx/src/main/scala/)

    • Provides GraphX-compatible APIs for migration
    • Bridges between DataFrame and RDD-based graph operations

Key Design Patterns

Spark Version Compatibility: Uses SparkShims pattern for version-specific implementations:

  • core/src/main/scala-spark-3/: Spark 3.x specific code
  • core/src/main/scala-spark-4/: Spark 4.x specific code
  • Common interface in main scala directory

DataFrame-Centric Design: Unlike GraphX's RDD approach, GraphFrames uses DataFrames throughout:

  • Vertices and edges are DataFrames with required schema (id column for vertices, src/dst for edges)
  • Leverages Spark SQL optimizations and Catalyst query planner
  • Allows mixing graph operations with relational queries

Algorithm Framework: Most algorithms follow this pattern:

  • Extend base traits like Logging, WithLocalCheckpoints, WithIntermediateStorageLevel
  • Use builder pattern for configuration (e.g., graph.pregel.withVertexColumn(...).sendMsgToDst(...).run())
  • Support checkpointing and persistence for iterative algorithms

Performance Considerations

Pregel Optimizations: The Pregel implementation includes automatic optimizations:

  • Detects when destination vertex state is unneeded and skips expensive joins
  • Conditional partitioning (source-only vs source+destination)
  • Automatic caching of intermediate results when beneficial

Memory Management:

  • Uses configurable storage levels for persistence
  • Automatic cleanup of intermediate DataFrames between iterations
  • Checkpoint support for long-running iterative algorithms

Testing Structure

Test Organization

  • Base test class: SparkFunSuite provides SparkContext setup
  • Algorithm tests: Each algorithm has corresponding *Suite.scala in lib/ subdirectory
  • Integration tests: GraphFrameSuite.scala for core API functionality
  • Pattern matching: PatternSuite.scala for motif finding features

Running Specific Tests

  • Algorithm-specific: build/sbt "testOnly *PageRankSuite"
  • Core functionality: build/sbt "testOnly *GraphFrameSuite"
  • Pregel framework: build/sbt "testOnly *PregelSuite"
  • All lib tests: build/sbt "testOnly org.graphframes.lib.*"

Multi-Version Testing

Tests run against multiple Spark versions in CI. The build system automatically:

  • Selects appropriate Scala versions based on Spark version
  • Uses version-specific SparkShims implementations
  • Validates compatibility across Spark 3.5+ and 4.0+

Multi-Language Support

Scala (Primary)

  • Main implementation in core/src/main/scala/
  • Uses Scala 2.12/2.13 depending on Spark version
  • Follows Spark's coding conventions and patterns

Python

  • Python bindings in python/graphframes/
  • Wraps Scala API using Spark's Python gateway
  • Separate test suite in python/tests/

Java

  • Java examples in core/src/main/java/
  • Uses Scala API through Java interop
  • Follows JavaBean conventions where applicable

Development Workflow

Code Quality

  • Formatting: scalafmt for Scala, black/isort/flake8 for Python
  • Style checking: scalafix for Scala linting
  • Pre-commit hooks: Available via pre-commit run all-files

Contribution Requirements

โš ๏ธ Legacy Codebase Rules:

  • NO breaking changes to existing APIs without explicit approval
  • NO modification of existing test behavior or CI configuration
  • NO large-scale refactoring or architectural changes
  • ALL changes must maintain backward compatibility

Standard Requirements:

  • All new algorithms must include comprehensive test coverage
  • Performance-critical code should include benchmarks (benchmarks/ directory)
  • Multi-version compatibility must be maintained across all supported Spark versions
  • Documentation updates required for new features
  • Zero test failures across all Spark versions before submitting PR

Release Process

The project uses automated releases through GitHub Actions:

  • Scala artifacts published to Maven Central
  • Python packages published to PyPI
  • Documentation deployed to graphframes.io

Repository README

Describes ItamarZand88/awesome-agent-conventions as a whole, which may contain artifacts other than this one. Where this artifact had no useful description of its own, its summary was taken from here.

Awesome Agent Conventions

A curated field guide to the convention files AI agents read, write, and act on.

22 conventions across 11 categories. From common project instruction files to newer agent-web discovery and trust formats.

Agent tools increasingly rely on plain files in a repository or website root: instructions, memory, rules, tool connections, prompt assets, discovery metadata, and protocol hints. The names are easy to mix up, and the adoption levels vary a lot.

This repo keeps the map practical:

  • Know what a file is for. Each entry names the convention, usual filename, primary readers, and spec or source.
  • Study real examples. Examples are fetched from public repositories by script, with provenance kept at the top of each file.
  • Separate practice from proposal. Maturity labels show what is widely used, what is early, and what is still only proposed.

Contents

What counts

This list is intentionally narrow. A file belongs here when it is a convention for agent behavior or agent-readable metadata: instructions, memory, skills, rules, tool config, prompt assets, or web-discovery hints.

Any file type can qualify - .md, .txt, .prompty, .json, dotfiles, or a directory pattern. Human-first project docs such as README.md, CONTRIBUTING.md, SECURITY.md, and CHANGELOG.md stay out unless the file has become an agent convention in its own right.

Maturity tiers

The badge is a claim about adoption, not quality. It keeps a proven convention from being presented the same way as a new idea.

BadgeTierMeaning
๐ŸŸขAdoptedUsed in production by multiple tools, projects, or teams.
๐ŸŸ EmergingPublished by a real organization, but still early or limited in adoption.
๐Ÿ”ตProposedPublicly described, but without clear adoption beyond the proposal.

Instruction & context

Standalone page: categories/instruction-context.md

ConventionFilesRead bySpec
๐ŸŸขAGENTS.mdAGENTS.mdMost coding agents - OpenAI Codex, Cursor, Jules, Aider, Gemini CLI, Zed, and othersspec โ†—
๐ŸŸขCLAUDE.mdCLAUDE.mdClaude Code, and tools that read the Claude memory conventionspec โ†—
๐ŸŸขTool-specific instruction filesGEMINI.md AGENT.md QWEN.md WARP.md CONVENTIONS.md copilot-instructions.mdEach file is read by its namesake tool - Gemini CLI, Amp, Qwen Code, Warp, Aider, GitHub Copilot - often alongside or as a bridge to AGENTS.mdspec โ†—
๐ŸŸ OKF (Open Knowledge Format).mdAgents over MCP (okfy, openknowledge, superops okf CLIs); Google's knowledge-catalog ingests bundlesspec โ†—
  • AGENTS.md - A plain-Markdown "README for agents" - build/test commands, conventions, and gotchas an agent needs before touching the code. The most widely adopted cross-tool instruction file.
  • CLAUDE.md - Anthropic's memory file for Claude Code - loaded automatically at session start to carry project commands, style rules, and standing instructions across turns.
  • Tool-specific instruction files - Per-tool instruction files that predate or coexist with AGENTS.md. Some tools now default to AGENTS.md while keeping legacy filenames alive, so these variants still matter when auditing real repositories.
  • OKF (Open Knowledge Format) - A machine-first organizational knowledge base: a version-controlled folder of typed Markdown files (one concept per file) that any agent reads as ground-truth context. Open-sourced by Google Cloud in 2026 as the content layer to MCP's transport.

Memory & state

Standalone page: categories/memory-state.md

ConventionFilesRead bySpec
๐ŸŸขMEMORY.mdMEMORY.mdClaude Code's auto-memory - the per-project MEMORY.md index it writes and re-reads each sessionspec โ†—
๐ŸŸขMemory Bankprojectbrief.md productContext.md activeContext.md systemPatterns.md techContext.md progress.mdCline, Roo Code, and Cursor (via the Memory Bank custom-instructions pattern)spec โ†—
  • MEMORY.md - A persistent, agent-maintained index of durable facts - written and re-read across sessions so an agent accumulates project memory instead of relearning each time.
  • Memory Bank - Cline's structured memory system - a set of Markdown files an agent reads at the start of every task to reconstruct full project context after its session memory resets. The six files shown are Cline's set; tools like Roo Code use an overlapping but different variant.

Spec-driven development

Standalone page: categories/spec-driven-development.md

ConventionFilesRead bySpec
๐ŸŸขSpec Kitconstitution.md spec.md plan.md tasks.mdGitHub Spec Kit's slash-command agents (Copilot, Claude, Gemini, Cursor, and more)spec โ†—
๐ŸŸขKiro steering filesproduct.md structure.md tech.mdAWS Kiro (steering files are largely Kiro-specific)spec โ†—
  • Spec Kit - GitHub's spec-driven workflow - a constitution plus per-feature spec โ†’ plan โ†’ tasks files that drive an agent through structured, reviewable implementation.
  • Kiro steering files - Kiro's always-on steering docs - product, structure, and tech files that give the agent persistent project context outside of any single spec.

Skills & prompt assets

Standalone page: categories/skills-prompt-assets.md

ConventionFilesRead bySpec
๐ŸŸขSKILL.mdSKILL.mdClaude Agent Skills, Claude Code, Amp, Agent Skills-compatible toolsspec โ†—
๐ŸŸขPrompt asset files.prompty .prompt system_prompt.txtPrompty tooling, Azure AI / Semantic Kernel, and apps that load externalized promptsspec โ†—
๐ŸŸขClaude Code commands.mdClaude Code - project .claude/commands/ and user ~/.claude/commands/spec โ†—
๐ŸŸขCopilot prompt & instruction files.prompt.md .instructions.mdGitHub Copilot in VS Code / Copilot CLIspec โ†—
  • SKILL.md - A self-contained, model-invoked capability file that tells an agent when to load a reusable procedure and how to execute it.
  • Prompt asset files - Externalized prompt files - Prompty's YAML-front-mattered .prompty, plain .prompt templates, and system_prompt.txt - that pull the prompt out of source code so it can be versioned and edited on its own. Only .prompty has a formal spec (prompty.ai); .prompt and system_prompt.txt are ad-hoc externalized-prompt filenames.
  • Claude Code commands - A Markdown file Claude Code exposes as a /slash-command - a reusable, version-controlled prompt workflow, with optional frontmatter (allowed-tools, model, argument-hint) and $ARGUMENTS and shell placeholders (@file references are a general Claude Code prompt feature, not command-specific). Now converging with Agent Skills, but still widely committed in its own right.
  • Copilot prompt & instruction files - Modular, path-scoped Copilot context: *.instructions.md auto-attach to matching files via an applyTo glob, while *.prompt.md are reusable prompts you invoke by name - the granular cousins of a single .github/copilot-instructions.md.

Tooling & connections

Standalone page: categories/tooling-connections.md

ConventionFilesRead bySpec
๐ŸŸขMCP server config.mcp.jsonClaude Code, Cursor, VS Code / Copilot, and Claude Desktop - every MCP host reads the same mcpServers schema, though the filename and path differ per toolspec โ†—
  • MCP server config - A JSON file that tells an agent which Model Context Protocol servers to launch and how (command, args, env) - making a project's tool and data integrations portable, shareable, and version-controlled across every MCP-capable client.

Rules & ignore files

Standalone page: categories/rules-ignore-files.md

ConventionFilesRead bySpec
๐ŸŸขRules files.cursorrules .mdc .clinerules .clinerules/ (pattern) .windsurfrulesCursor (.cursorrules / .mdc), Cline (.clinerules/ and legacy .clinerules), Windsurf (.windsurfrules)spec โ†—
๐ŸŸขAI ignore files.aiignore .cursorignore .codeiumignore .aiexcludeJetBrains Junie (.aiignore), Cursor (.cursorignore), Codeium/Windsurf (.codeiumignore)spec โ†—
  • Rules files - Per-tool rule files that scope agent behavior - older single-file forms (.cursorrules, .clinerules, .windsurfrules) and newer directory-based, glob-scoped forms (.cursor/rules/.mdc, .clinerules/, .windsurf/rules/.md).
  • AI ignore files - gitignore-syntax files that fence an AI agent out of paths - secrets, vendored code, generated output - so they're never sent to the model as context.

Design

Standalone page: categories/design.md

ConventionFilesRead bySpec
๐ŸŸขDESIGN.mdDESIGN.mdGoogle Stitch natively; and coding agents (e.g. Claude Code) when pointed at it as design contextspec โ†—
  • DESIGN.md - A structured, machine-readable design specification - tokens, components, and layout intent - that an agent reads to generate or keep UI consistent with an established system. Open-sourced by Google Labs in 2026 as a cross-tool draft spec.

Web & discoverability

Standalone page: categories/web-discoverability.md

ConventionFilesRead bySpec
๐ŸŸขllms.txtllms.txt llms-full.txt (pattern)Docs sites publish it for LLM tools and crawlers - though no major provider has confirmed reading itspec โ†—
๐ŸŸขpricing.mdpricing.mdAgents and LLM browsers fetching a clean, parse-able pricing pagespec โ†—
  • llms.txt - A proposed-turned-widely-published standard: a root-level Markdown file giving LLMs a curated, link-rich map of a site's docs. Published across hundreds of developer-docs sites - though whether the major LLM providers actually read it remains unproven.
  • pricing.md - The Markdown twin of a pricing page - same URL with a .md suffix - so an agent gets structured plans and numbers instead of scraping marketing HTML. A concrete, shipping instance of the page.md pattern.

Agent-web trust

Standalone page: categories/agent-web-trust.md

ConventionFilesRead bySpec
๐ŸŸ auth.mdauth.mdAgents discovering how to authenticate to a service (early adopters)spec โ†—
๐Ÿ”ตai.txtai.txtAI training/data-mining crawlers that voluntarily honor AI usage preferences; crawler support is not yet reliablespec โ†—
  • auth.md - A Markdown file that tells an agent how to authenticate with a service - discovery of auth endpoints and flows. Shipped by WorkOS as a real, working convention, but adoption beyond it is still early.
  • ai.txt - A text file declaring machine-readable consent, licensing, or policy preferences for AI training and data-mining. Spawning popularized the deployed root-file pattern, and a 2026 Internet-Draft now proposes a well-known URI; adoption and crawler obedience are still thin, so it stays ๐Ÿ”ต.

Identity & protocols

Standalone page: categories/identity-protocols.md

ConventionFilesRead bySpec
๐ŸŸ Agent Cards (A2A)agent-card.json agent.json (pattern)A2A-compatible agents discovering another agent's capabilitiesspec โ†—
  • Agent Cards (A2A) - The Agent2Agent (A2A) capability card - a JSON document at a well-known path advertising an agent's skills, endpoints, and auth so other agents can discover and call it. Now a Linux Foundation project at v1.0; adoption is growing but early.

Proposed namespace

Standalone page: categories/proposed-namespace.md

ConventionFilesRead bySpec
๐Ÿ”ตThe protocols.md namespaceproof.md- (no demonstrated readers; aspirational)spec โ†—
  • The protocols.md namespace - A single maintainer's pre-registered namespace of ~74 aspirational .md "protocols" (proof.md, signature.md, reputation.md, โ€ฆ) staked as Schelling points for a future agent web. Published concept, no demonstrated adoption - see the page for the audited, honest caveats.

Maintaining examples

Example files are fetched, not invented. The extractor pulls them from public sources, stores them under conventions/<slug>/examples/<source>/<filename>, and adds a line-1 provenance comment. The examples remain under their upstream owners' licenses and terms; see THIRD_PARTY_EXAMPLES.md before reusing them.

To refresh everything:

pip install -r scripts/requirements.txt
python scripts/extract.py          # fetch real files + rebuild each convention's README
python scripts/build_readme.py     # rebuild this README from scripts/targets.json

Re-running is idempotent. A missing target prints a miss and is skipped. Examples are representative samples: any file over 256 KB (for example, a multi-MB llms-full.txt) is truncated with a marker pointing back to the full source. scripts/targets.json remains the source of truth for conventions that have not been migrated yet; the skill-md pilot uses local convention metadata instead. Edit the relevant source and re-run both scripts.

Shortcut targets are available in the Makefile:

make verify          # schema + generated files + example provenance + links
make extract         # refetch public examples and rebuild generated docs
make license-report  # summarize upstream licenses for vendored examples

CI keeps the generated files and links honest. The verify workflow checks that generated docs match catalog metadata and migrated local metadata, and that every spec, example, and instance URL still resolves on each pull request and weekly. Run the same link check locally with python scripts/check_links.py.

Contributing

Read CONTRIBUTING.md. In short: an entry must pass the filter above and carry evidence for its maturity tier. Add sources to scripts/targets.json for non-migrated conventions, run the scripts, and open a PR. The skill-md pilot uses local convention metadata instead. Do not hand-write example files.

Before proposing adjacent standards, check WATCHLIST.md. Project direction lives in ROADMAP.md.

License

The curation, scripts, and original prose in this repository are MIT. Vendored example files remain under their upstream owners' licenses and terms; see THIRD_PARTY_EXAMPLES.md.

Trustgrade A

  • passBody integrity

    Whether the stored document is plausibly the kind of file the artifact declares, rather than something fetched by mistake.

  • passType matchnot applicable to this artifact type

    Whether the artifact is really the kind of thing its metadata claims it is.

  • passFreshness

    How long since the source repository was last pushed to.

  • passPrompt injection

    Scans the artifact's own text for instructions aimed at your agent rather than at you.

  • passLicense

    Whether the source repository declares an SPDX license permissive enough to redistribute.

How the grade is calculated

Each check contributes 0 points when it passes, 1 when it warns, and 2 when it fails. The total maps to a letter:

  • Aevery check passed
  • Bone warning
  • Ctwo warnings
  • Dprompt injection or body integrity failed, or three warnings
  • Fone of those failed, and something else is wrong

These are automated hygiene checks, not a security audit, and not a dependency or vulnerability scan. A grade of A means nothing was flagged โ€” not that the artifact is safe.

Versions

  • git-a022a8588f982026-08-04