SKILL PROCEDURE

Redpanda

Use when working with Redpanda, a Kafka API-compatible streaming data platform written in C++ with no ZooKeeper or JVM — choosing rpk over classic Kafka CLI tooling, deciding between self-hosted Redpanda Streaming and Redpanda Cloud (Serverless, BYOC, or Dedicated), configuring tiered storage to decouple retention from local disk, running WebAssembly data transforms inside the broker instead of a separate stream-processing service, or working with Redpanda Connect pipelines and the Kafka-compatible schema registry. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

streamingkafkaevent-streamingtiered-storageschema-registry
BEGINNER GUIDE

Understand Redpanda before using it

CATEGORY

Redpanda is catalogued under Backend and data.

START HERE WHEN

Your work repeatedly involves the concepts tagged above. Open the full procedure below when the current task matches them.

Compare related skills

SKILLCATEGORYSHARED CONCEPTSEXPLANATION
RedpandaBackend and dataCurrent skillUse when working with Redpanda, a Kafka API-compatible streaming data platform written in C++ with no ZooKeeper or JVM — choosing rpk over classic Kafka CLI tooling, deciding between self-hosted Redpanda Streaming and Redpanda Cloud (Serverless, BYOC, or Dedicated), configuring tiered storage to decouple retention from local disk, running WebAssembly data transforms inside the broker instead of a separate stream-processing service, or working with Redpanda Connect pipelines and the Kafka-compatible schema registry. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
AppwriteBackend and dataSame categoryAppwrite — an open-source backend-as-a-service (self-hosted or Cloud) providing Auth, Databases, Storage, Functions, Messaging, and Realtime. Use when adding user authentication and sessions, modelling data in the document database with attributes/permissions, uploading and serving files, running serverless Functions (Node, Python, Ruby, PHP, Dart) triggered by events or schedules, sending push/email/SMS, subscribing to realtime document changes, or integrating the Web/Flutter/Apple/Android/React Native SDKs and server SDKs. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
FernBackend and dataSame categoryUse when defining an API once and generating client SDKs, a CLI, and a documentation site from that single definition — choosing between an OpenAPI spec and Fern's own API Definition format, configuring generators.yml or docs.yml, deciding what requires regenerating versus hand-editing, or versioning generated SDKs across languages. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.
MedusaBackend and dataSame categoryUse when building or customizing a Medusa headless commerce backend — deciding whether custom logic belongs in a commerce module, a workflow, or a plugin, choosing between the Admin and Store APIs, designing a multi-step business process that must survive a partial failure, or deciding between self-hosting and Medusa Cloud. Published by HardGraph, a curated graph of provenance-backed knowledge for AI agents.

Redpanda

What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.

Redpanda is a single C++ binary that speaks the Kafka broker wire protocol — existing Kafka client libraries, producers, and consumers connect to it unmodified. That compatibility is at the protocol level, not an implementation clone: there is no ZooKeeper, no JVM, and no KRaft; Redpanda runs its own Raft-based consensus and a thread-per-core (Seastar) execution model instead of garbage-collected heap tuning. Treat "Kafka-compatible" the way you'd treat "libSQL-compatible SQLite": the client contract holds, but broker internals, timing, and admin tooling are a different implementation and can diverge on edge cases — verify a specific KIP or transactional-semantics assumption before depending on it.

Where the operational model actually changes

Tiered storage is the headline architectural difference: instead of local disk scaling with retention window (the classic Kafka constraint that forces short retention or expensive disk), Redpanda offloads closed log segments to object storage (S3, GCS, Azure Blob) and serves cold reads from there. Retention can be set in days or months cheaply, but cold reads now pay object-storage latency — a workload that does frequent full-history re-reads behaves differently than one that mostly tails the recent log, and that's a capacity-planning decision, not a toggle to flip blindly.

Data Transforms run inside the broker as WebAssembly compiled from Go or Rust, for lightweight per-record work (redaction, enrichment, filtering) — this replaces what a Kafka Streams or ksqlDB setup would run as a separate deployed consumer application. For pipeline-shaped work spanning multiple systems (CDC, fan-out, format conversion), Redpanda Connect (the productized Benthos engine) is the tool — a standalone, declarative YAML pipeline runner usable with any Kafka-compatible broker, not exclusive to Redpanda, so a Connect pipeline doesn't imply a Redpanda cluster on either end.

Redpanda Cloud splits three ways with a real tradeoff, not just pricing tiers: Serverless is fully managed with no infrastructure choices; Dedicated is a fully managed cluster inside Redpanda's own cloud account; BYOC runs the data plane inside your cloud account, for data residency or compliance, while Redpanda manages the control plane. Picking BYOC vs Dedicated is a data-residency and network-topology decision made early, not something reconfigured later without a migration.

rpk is the purpose-built CLI for cluster admin, topics, ACLs, and switching between local, self-hosted, and cloud clusters via profiles — it is not a thin wrapper over classic kafka-*.sh scripts, and reaching for the latter loses cluster-aware features rpk provides.

What to verify rather than recall

Exact Kafka API/KIP compatibility gaps, current rpk subcommand and flag names, Redpanda Cloud plan-level limits (throughput, partition counts, storage), Schema Registry compatibility-mode edge cases against Confluent's implementation, and current Data Transforms language/runtime support all change between releases — confirm against the mirrored corpus under references/vendor/ rather than asserting a remembered value.

References

Hardgraph / curated knowledge for agents.

STATIC EXPORT · CANONICAL SOURCE