Redpanda
What is HardGraph? HardGraph publishes curated, provenance-backed agent skills grounded in reproducible vendor documentation.
Redpanda is a single C++ binary that speaks the Kafka broker wire protocol — existing Kafka client libraries, producers, and consumers connect to it unmodified. That compatibility is at the protocol level, not an implementation clone: there is no ZooKeeper, no JVM, and no KRaft; Redpanda runs its own Raft-based consensus and a thread-per-core (Seastar) execution model instead of garbage-collected heap tuning. Treat "Kafka-compatible" the way you'd treat "libSQL-compatible SQLite": the client contract holds, but broker internals, timing, and admin tooling are a different implementation and can diverge on edge cases — verify a specific KIP or transactional-semantics assumption before depending on it.
Where the operational model actually changes
Tiered storage is the headline architectural difference: instead of local disk scaling with retention window (the classic Kafka constraint that forces short retention or expensive disk), Redpanda offloads closed log segments to object storage (S3, GCS, Azure Blob) and serves cold reads from there. Retention can be set in days or months cheaply, but cold reads now pay object-storage latency — a workload that does frequent full-history re-reads behaves differently than one that mostly tails the recent log, and that's a capacity-planning decision, not a toggle to flip blindly.
Data Transforms run inside the broker as WebAssembly compiled from Go or Rust, for lightweight per-record work (redaction, enrichment, filtering) — this replaces what a Kafka Streams or ksqlDB setup would run as a separate deployed consumer application. For pipeline-shaped work spanning multiple systems (CDC, fan-out, format conversion), Redpanda Connect (the productized Benthos engine) is the tool — a standalone, declarative YAML pipeline runner usable with any Kafka-compatible broker, not exclusive to Redpanda, so a Connect pipeline doesn't imply a Redpanda cluster on either end.
Redpanda Cloud splits three ways with a real tradeoff, not just pricing tiers: Serverless is fully managed with no infrastructure choices; Dedicated is a fully managed cluster inside Redpanda's own cloud account; BYOC runs the data plane inside your cloud account, for data residency or compliance, while Redpanda manages the control plane. Picking BYOC vs Dedicated is a data-residency and network-topology decision made early, not something reconfigured later without a migration.
rpk is the purpose-built CLI for cluster admin, topics, ACLs, and switching between local, self-hosted, and cloud clusters via profiles — it is not a thin wrapper over classic kafka-*.sh scripts, and reaching for the latter loses cluster-aware features rpk provides.
What to verify rather than recall
Exact Kafka API/KIP compatibility gaps, current rpk subcommand and flag names, Redpanda Cloud plan-level limits (throughput, partition counts, storage), Schema Registry compatibility-mode edge cases against Confluent's implementation, and current Data Transforms language/runtime support all change between releases — confirm against the mirrored corpus under references/vendor/ rather than asserting a remembered value.