@itamarzand88/awesome-agent-conventions-12
AA curated field guide to the convention files AI agents read, write, and act on.
Install
agr install @itamarzand88/awesome-agent-conventions-12 --target codexWrites 1 file into AGENTS.md, pinned to git-a022a858.
- AGENTS.md
Document
AGENTS.md
This file provides guidance to AI agents when working with code in this repository.
๐จ Legacy Codebase Guidelines ๐จ
CRITICAL: This is a mature, mission-critical codebase (10+ years old in some areas). Follow these rules strictly:
Regression Prevention
- NEVER introduce breaking changes without explicit approval
- NEVER modify existing public APIs unless specifically requested
- NEVER change existing test assertions or remove tests without approval
- NEVER modify CI/build configuration without explicit request
- ALWAYS run the complete test suite before committing changes
- ALWAYS test against all supported Spark versions
Code Change Philosophy
- Prefer addition over modification: Add new functionality alongside existing code
- Preserve existing behavior: If unsure, err on the side of maintaining status quo
- No "cowboy-style" refactoring: Avoid large-scale structural changes
- Conservative by default: Only make changes that directly address the specific request
When in Doubt
- Ask before changing: If a change seems risky, request explicit approval
- Start small: Make minimal changes to achieve the goal
- Maintain backward compatibility: Existing user code must continue to work
- Document reasoning: Explain why changes are necessary in commit messages
Build and Development Commands
Core Build System
- Build tool: sbt (Scala Build Tool) version 1.11.0
- Build project:
build/sbt compile - Run all tests:
build/sbt test - Run single test suite:
build/sbt "testOnly *PregelSuite" - Run single test:
build/sbt "testOnly *PregelSuite -- -z 'test name pattern'" - Check code formatting:
build/sbt scalafmtCheckAll - Apply code formatting:
build/sbt scalafmtAll - Check code style:
build/sbt "scalafixAll --check" - Generate documentation:
build/sbt doc - Run with coverage:
build/sbt "coverage test coverageReport"
Multi-Spark Version Testing
The project supports multiple Spark versions. Use -Dspark.version=X.Y.Z to specify:
build/sbt -Dspark.version=3.5.7 testbuild/sbt -Dspark.version=4.0.1 testbuild/sbt -Dspark.version=4.1.0 test
Pre-PR Checklist Commands
โ ๏ธ MANDATORY: These commands MUST pass before any PR. Failures indicate potential regressions.
Before raising a pull request, run these commands to ensure code quality and compatibility:
- Format code:
build/sbt scalafmtAll - Check formatting:
build/sbt scalafmtCheckAll - Check style/linting:
build/sbt "scalafixAll --check" - Build documentation:
build/sbt doc - Run tests with coverage:
build/sbt "coverage test coverageReport" - Test against multiple Spark versions:
build/sbt -Dspark.version=3.5.7 testbuild/sbt -Dspark.version=4.0.1 testbuild/sbt -Dspark.version=4.1.0 test
Alternative using pre-commit: pre-commit run all-files (handles Python formatting: black, isort, flake8)
Python Components
- Python package location:
/python/directory - Python tests: Located in
/python/tests/
Architecture Overview
Core Structure
GraphFrames is a graph processing library built on Apache Spark DataFrames with three main components:
-
Core GraphFrame API (
core/src/main/scala/org/graphframes/)GraphFrame.scala: Main graph abstraction using DataFrames for vertices/edges- Provides graph algorithms, motif finding, and subgraph operations
-
Algorithm Library (
core/src/main/scala/org/graphframes/lib/)- Standard graph algorithms: PageRank, ConnectedComponents, ShortestPaths, etc.
Pregel.scala: Implements Pregel-style bulk synchronous parallel processingAggregateMessages.scala: Lower-level message-passing API
-
GraphX Compatibility Layer (
graphx/src/main/scala/)- Provides GraphX-compatible APIs for migration
- Bridges between DataFrame and RDD-based graph operations
Key Design Patterns
Spark Version Compatibility: Uses SparkShims pattern for version-specific implementations:
core/src/main/scala-spark-3/: Spark 3.x specific codecore/src/main/scala-spark-4/: Spark 4.x specific code- Common interface in main scala directory
DataFrame-Centric Design: Unlike GraphX's RDD approach, GraphFrames uses DataFrames throughout:
- Vertices and edges are DataFrames with required schema (id column for vertices, src/dst for edges)
- Leverages Spark SQL optimizations and Catalyst query planner
- Allows mixing graph operations with relational queries
Algorithm Framework: Most algorithms follow this pattern:
- Extend base traits like
Logging,WithLocalCheckpoints,WithIntermediateStorageLevel - Use builder pattern for configuration (e.g.,
graph.pregel.withVertexColumn(...).sendMsgToDst(...).run()) - Support checkpointing and persistence for iterative algorithms
Performance Considerations
Pregel Optimizations: The Pregel implementation includes automatic optimizations:
- Detects when destination vertex state is unneeded and skips expensive joins
- Conditional partitioning (source-only vs source+destination)
- Automatic caching of intermediate results when beneficial
Memory Management:
- Uses configurable storage levels for persistence
- Automatic cleanup of intermediate DataFrames between iterations
- Checkpoint support for long-running iterative algorithms
Testing Structure
Test Organization
- Base test class:
SparkFunSuiteprovides SparkContext setup - Algorithm tests: Each algorithm has corresponding
*Suite.scalainlib/subdirectory - Integration tests:
GraphFrameSuite.scalafor core API functionality - Pattern matching:
PatternSuite.scalafor motif finding features
Running Specific Tests
- Algorithm-specific:
build/sbt "testOnly *PageRankSuite" - Core functionality:
build/sbt "testOnly *GraphFrameSuite" - Pregel framework:
build/sbt "testOnly *PregelSuite" - All lib tests:
build/sbt "testOnly org.graphframes.lib.*"
Multi-Version Testing
Tests run against multiple Spark versions in CI. The build system automatically:
- Selects appropriate Scala versions based on Spark version
- Uses version-specific SparkShims implementations
- Validates compatibility across Spark 3.5+ and 4.0+
Multi-Language Support
Scala (Primary)
- Main implementation in
core/src/main/scala/ - Uses Scala 2.12/2.13 depending on Spark version
- Follows Spark's coding conventions and patterns
Python
- Python bindings in
python/graphframes/ - Wraps Scala API using Spark's Python gateway
- Separate test suite in
python/tests/
Java
- Java examples in
core/src/main/java/ - Uses Scala API through Java interop
- Follows JavaBean conventions where applicable
Development Workflow
Code Quality
- Formatting: scalafmt for Scala, black/isort/flake8 for Python
- Style checking: scalafix for Scala linting
- Pre-commit hooks: Available via
pre-commit run all-files
Contribution Requirements
โ ๏ธ Legacy Codebase Rules:
- NO breaking changes to existing APIs without explicit approval
- NO modification of existing test behavior or CI configuration
- NO large-scale refactoring or architectural changes
- ALL changes must maintain backward compatibility
Standard Requirements:
- All new algorithms must include comprehensive test coverage
- Performance-critical code should include benchmarks (
benchmarks/directory) - Multi-version compatibility must be maintained across all supported Spark versions
- Documentation updates required for new features
- Zero test failures across all Spark versions before submitting PR
Release Process
The project uses automated releases through GitHub Actions:
- Scala artifacts published to Maven Central
- Python packages published to PyPI
- Documentation deployed to graphframes.io
Repository README
Describes ItamarZand88/awesome-agent-conventions as a whole, which may contain artifacts other than this one. Where this artifact had no useful description of its own, its summary was taken from here.
Awesome Agent Conventions
A curated field guide to the convention files AI agents read, write, and act on.
22 conventions across 11 categories. From common project instruction files to newer agent-web discovery and trust formats.
Agent tools increasingly rely on plain files in a repository or website root: instructions, memory, rules, tool connections, prompt assets, discovery metadata, and protocol hints. The names are easy to mix up, and the adoption levels vary a lot.
This repo keeps the map practical:
- Know what a file is for. Each entry names the convention, usual filename, primary readers, and spec or source.
- Study real examples. Examples are fetched from public repositories by script, with provenance kept at the top of each file.
- Separate practice from proposal. Maturity labels show what is widely used, what is early, and what is still only proposed.
Contents
- Instruction & context ยท category page
- ๐ข AGENTS.md
- ๐ข CLAUDE.md
- ๐ข Tool-specific instruction files
- ๐ OKF (Open Knowledge Format)
- Memory & state ยท category page
- ๐ข MEMORY.md
- ๐ข Memory Bank
- Spec-driven development ยท category page
- ๐ข Spec Kit
- ๐ข Kiro steering files
- Skills & prompt assets ยท category page
- ๐ข SKILL.md
- ๐ข Prompt asset files
- ๐ข Claude Code commands
- ๐ข Copilot prompt & instruction files
- Tooling & connections ยท category page
- ๐ข MCP server config
- Rules & ignore files ยท category page
- ๐ข Rules files
- ๐ข AI ignore files
- Design ยท category page
- ๐ข DESIGN.md
- Web & discoverability ยท category page
- ๐ข llms.txt
- ๐ข pricing.md
- Agent-web trust ยท category page
- Identity & protocols ยท category page
- Proposed namespace ยท category page
What counts
This list is intentionally narrow. A file belongs here when it is a convention for agent behavior or agent-readable metadata: instructions, memory, skills, rules, tool config, prompt assets, or web-discovery hints.
Any file type can qualify - .md, .txt, .prompty, .json, dotfiles, or a
directory pattern. Human-first project docs such as README.md,
CONTRIBUTING.md, SECURITY.md, and CHANGELOG.md stay out unless the file has
become an agent convention in its own right.
Maturity tiers
The badge is a claim about adoption, not quality. It keeps a proven convention from being presented the same way as a new idea.
| Badge | Tier | Meaning |
|---|---|---|
| ๐ข | Adopted | Used in production by multiple tools, projects, or teams. |
| ๐ | Emerging | Published by a real organization, but still early or limited in adoption. |
| ๐ต | Proposed | Publicly described, but without clear adoption beyond the proposal. |
Instruction & context
Standalone page: categories/instruction-context.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | AGENTS.md | AGENTS.md | Most coding agents - OpenAI Codex, Cursor, Jules, Aider, Gemini CLI, Zed, and others | spec โ |
| ๐ข | CLAUDE.md | CLAUDE.md | Claude Code, and tools that read the Claude memory convention | spec โ |
| ๐ข | Tool-specific instruction files | GEMINI.md AGENT.md QWEN.md WARP.md CONVENTIONS.md copilot-instructions.md | Each file is read by its namesake tool - Gemini CLI, Amp, Qwen Code, Warp, Aider, GitHub Copilot - often alongside or as a bridge to AGENTS.md | spec โ |
| ๐ | OKF (Open Knowledge Format) | .md | Agents over MCP (okfy, openknowledge, superops okf CLIs); Google's knowledge-catalog ingests bundles | spec โ |
- AGENTS.md - A plain-Markdown "README for agents" - build/test commands, conventions, and gotchas an agent needs before touching the code. The most widely adopted cross-tool instruction file.
- CLAUDE.md - Anthropic's memory file for Claude Code - loaded automatically at session start to carry project commands, style rules, and standing instructions across turns.
- Tool-specific instruction files - Per-tool instruction files that predate or coexist with AGENTS.md. Some tools now default to AGENTS.md while keeping legacy filenames alive, so these variants still matter when auditing real repositories.
- OKF (Open Knowledge Format) - A machine-first organizational knowledge base: a version-controlled folder of typed Markdown files (one concept per file) that any agent reads as ground-truth context. Open-sourced by Google Cloud in 2026 as the content layer to MCP's transport.
Memory & state
Standalone page: categories/memory-state.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | MEMORY.md | MEMORY.md | Claude Code's auto-memory - the per-project MEMORY.md index it writes and re-reads each session | spec โ |
| ๐ข | Memory Bank | projectbrief.md productContext.md activeContext.md systemPatterns.md techContext.md progress.md | Cline, Roo Code, and Cursor (via the Memory Bank custom-instructions pattern) | spec โ |
- MEMORY.md - A persistent, agent-maintained index of durable facts - written and re-read across sessions so an agent accumulates project memory instead of relearning each time.
- Memory Bank - Cline's structured memory system - a set of Markdown files an agent reads at the start of every task to reconstruct full project context after its session memory resets. The six files shown are Cline's set; tools like Roo Code use an overlapping but different variant.
Spec-driven development
Standalone page: categories/spec-driven-development.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | Spec Kit | constitution.md spec.md plan.md tasks.md | GitHub Spec Kit's slash-command agents (Copilot, Claude, Gemini, Cursor, and more) | spec โ |
| ๐ข | Kiro steering files | product.md structure.md tech.md | AWS Kiro (steering files are largely Kiro-specific) | spec โ |
- Spec Kit - GitHub's spec-driven workflow - a constitution plus per-feature spec โ plan โ tasks files that drive an agent through structured, reviewable implementation.
- Kiro steering files - Kiro's always-on steering docs - product, structure, and tech files that give the agent persistent project context outside of any single spec.
Skills & prompt assets
Standalone page: categories/skills-prompt-assets.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | SKILL.md | SKILL.md | Claude Agent Skills, Claude Code, Amp, Agent Skills-compatible tools | spec โ |
| ๐ข | Prompt asset files | .prompty .prompt system_prompt.txt | Prompty tooling, Azure AI / Semantic Kernel, and apps that load externalized prompts | spec โ |
| ๐ข | Claude Code commands | .md | Claude Code - project .claude/commands/ and user ~/.claude/commands/ | spec โ |
| ๐ข | Copilot prompt & instruction files | .prompt.md .instructions.md | GitHub Copilot in VS Code / Copilot CLI | spec โ |
- SKILL.md - A self-contained, model-invoked capability file that tells an agent when to load a reusable procedure and how to execute it.
- Prompt asset files - Externalized prompt files - Prompty's YAML-front-mattered .prompty, plain .prompt templates, and system_prompt.txt - that pull the prompt out of source code so it can be versioned and edited on its own. Only .prompty has a formal spec (prompty.ai); .prompt and system_prompt.txt are ad-hoc externalized-prompt filenames.
- Claude Code commands - A Markdown file Claude Code exposes as a /slash-command - a reusable, version-controlled prompt workflow, with optional frontmatter (allowed-tools, model, argument-hint) and $ARGUMENTS and shell placeholders (@file references are a general Claude Code prompt feature, not command-specific). Now converging with Agent Skills, but still widely committed in its own right.
- Copilot prompt & instruction files - Modular, path-scoped Copilot context: *.instructions.md auto-attach to matching files via an applyTo glob, while *.prompt.md are reusable prompts you invoke by name - the granular cousins of a single .github/copilot-instructions.md.
Tooling & connections
Standalone page: categories/tooling-connections.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | MCP server config | .mcp.json | Claude Code, Cursor, VS Code / Copilot, and Claude Desktop - every MCP host reads the same mcpServers schema, though the filename and path differ per tool | spec โ |
- MCP server config - A JSON file that tells an agent which Model Context Protocol servers to launch and how (command, args, env) - making a project's tool and data integrations portable, shareable, and version-controlled across every MCP-capable client.
Rules & ignore files
Standalone page: categories/rules-ignore-files.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | Rules files | .cursorrules .mdc .clinerules .clinerules/ (pattern) .windsurfrules | Cursor (.cursorrules / .mdc), Cline (.clinerules/ and legacy .clinerules), Windsurf (.windsurfrules) | spec โ |
| ๐ข | AI ignore files | .aiignore .cursorignore .codeiumignore .aiexclude | JetBrains Junie (.aiignore), Cursor (.cursorignore), Codeium/Windsurf (.codeiumignore) | spec โ |
- Rules files - Per-tool rule files that scope agent behavior - older single-file forms (.cursorrules, .clinerules, .windsurfrules) and newer directory-based, glob-scoped forms (.cursor/rules/.mdc, .clinerules/, .windsurf/rules/.md).
- AI ignore files - gitignore-syntax files that fence an AI agent out of paths - secrets, vendored code, generated output - so they're never sent to the model as context.
Design
Standalone page: categories/design.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | DESIGN.md | DESIGN.md | Google Stitch natively; and coding agents (e.g. Claude Code) when pointed at it as design context | spec โ |
- DESIGN.md - A structured, machine-readable design specification - tokens, components, and layout intent - that an agent reads to generate or keep UI consistent with an established system. Open-sourced by Google Labs in 2026 as a cross-tool draft spec.
Web & discoverability
Standalone page: categories/web-discoverability.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ข | llms.txt | llms.txt llms-full.txt (pattern) | Docs sites publish it for LLM tools and crawlers - though no major provider has confirmed reading it | spec โ |
| ๐ข | pricing.md | pricing.md | Agents and LLM browsers fetching a clean, parse-able pricing page | spec โ |
- llms.txt - A proposed-turned-widely-published standard: a root-level Markdown file giving LLMs a curated, link-rich map of a site's docs. Published across hundreds of developer-docs sites - though whether the major LLM providers actually read it remains unproven.
- pricing.md - The Markdown twin of a pricing page - same URL with a .md suffix - so an agent gets structured plans and numbers instead of scraping marketing HTML. A concrete, shipping instance of the page.md pattern.
Agent-web trust
Standalone page: categories/agent-web-trust.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ | auth.md | auth.md | Agents discovering how to authenticate to a service (early adopters) | spec โ |
| ๐ต | ai.txt | ai.txt | AI training/data-mining crawlers that voluntarily honor AI usage preferences; crawler support is not yet reliable | spec โ |
- auth.md - A Markdown file that tells an agent how to authenticate with a service - discovery of auth endpoints and flows. Shipped by WorkOS as a real, working convention, but adoption beyond it is still early.
- ai.txt - A text file declaring machine-readable consent, licensing, or policy preferences for AI training and data-mining. Spawning popularized the deployed root-file pattern, and a 2026 Internet-Draft now proposes a well-known URI; adoption and crawler obedience are still thin, so it stays ๐ต.
Identity & protocols
Standalone page: categories/identity-protocols.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ | Agent Cards (A2A) | agent-card.json agent.json (pattern) | A2A-compatible agents discovering another agent's capabilities | spec โ |
- Agent Cards (A2A) - The Agent2Agent (A2A) capability card - a JSON document at a well-known path advertising an agent's skills, endpoints, and auth so other agents can discover and call it. Now a Linux Foundation project at v1.0; adoption is growing but early.
Proposed namespace
Standalone page: categories/proposed-namespace.md
| Convention | Files | Read by | Spec | |
|---|---|---|---|---|
| ๐ต | The protocols.md namespace | proof.md | - (no demonstrated readers; aspirational) | spec โ |
- The protocols.md namespace - A single maintainer's pre-registered namespace of ~74 aspirational .md "protocols" (proof.md, signature.md, reputation.md, โฆ) staked as Schelling points for a future agent web. Published concept, no demonstrated adoption - see the page for the audited, honest caveats.
Maintaining examples
Example files are fetched, not invented. The extractor pulls them from public
sources, stores them under conventions/<slug>/examples/<source>/<filename>,
and adds a line-1 provenance comment. The examples remain under their upstream
owners' licenses and terms; see
THIRD_PARTY_EXAMPLES.md before reusing them.
To refresh everything:
pip install -r scripts/requirements.txt
python scripts/extract.py # fetch real files + rebuild each convention's README
python scripts/build_readme.py # rebuild this README from scripts/targets.json
Re-running is idempotent. A missing target prints a miss and is skipped.
Examples are representative samples: any file over 256 KB (for example, a
multi-MB llms-full.txt) is truncated with a marker pointing back to the full
source. scripts/targets.json remains the source of truth for conventions that
have not been migrated yet; the skill-md pilot uses local convention metadata
instead. Edit the relevant source and re-run both scripts.
Shortcut targets are available in the Makefile:
make verify # schema + generated files + example provenance + links
make extract # refetch public examples and rebuild generated docs
make license-report # summarize upstream licenses for vendored examples
CI keeps the generated files and links honest. The
verify workflow checks that generated docs match
catalog metadata and migrated local metadata, and that every spec, example, and
instance URL still resolves on each pull request and weekly. Run the same link
check locally with
python scripts/check_links.py.
Contributing
Read CONTRIBUTING.md. In short: an entry must pass the filter
above and carry evidence for its maturity tier. Add sources to
scripts/targets.json for non-migrated conventions, run the scripts, and open
a PR. The skill-md pilot uses local convention metadata instead. Do not
hand-write example files.
Before proposing adjacent standards, check WATCHLIST.md. Project direction lives in ROADMAP.md.
License
The curation, scripts, and original prose in this repository are MIT. Vendored example files remain under their upstream owners' licenses and terms; see THIRD_PARTY_EXAMPLES.md.
Trustgrade A
- passBody integrity
Whether the stored document is plausibly the kind of file the artifact declares, rather than something fetched by mistake.
- passType matchnot applicable to this artifact type
Whether the artifact is really the kind of thing its metadata claims it is.
- passFreshness
How long since the source repository was last pushed to.
- passPrompt injection
Scans the artifact's own text for instructions aimed at your agent rather than at you.
- passLicense
Whether the source repository declares an SPDX license permissive enough to redistribute.
How the grade is calculated
Each check contributes 0 points when it passes, 1 when it warns, and 2 when it fails. The total maps to a letter:
- Aevery check passed
- Bone warning
- Ctwo warnings
- Dprompt injection or body integrity failed, or three warnings
- Fone of those failed, and something else is wrong
These are automated hygiene checks, not a security audit, and not a dependency or vulnerability scan. A grade of A means nothing was flagged โ not that the artifact is safe.
Versions
git-a022a8588f982026-08-04