TL;DR: Learn a compact set of CLI commands across Git, Docker, kubectl, Terraform, and your cloud provider CLI; prioritize CI/CD pipeline patterns, container orchestration, and repeatable IaC module scaffolds. Examples below link to a ready repo with common commands and scaffolds: DevOps commands.
Core DevOps CLI Commands: what to master and why
Being fluent with a small set of command-line tools yields outsized productivity. At the center: Git (source control), Docker and container runtimes (build, run, inspect), kubectl (interact with Kubernetes), and your chosen IaC CLI (terraform, az cli, aws cli, gcloud). Each tool provides a handful of commands you’ll use daily — commit and branch management in Git, image builds and container inspect in Docker, apply/get/logs in kubectl, and plan/apply/state in Terraform.
Start by learning safe, repeatable patterns: use feature branches (git checkout -b), automated builds (docker build), and declarative deployment (kubectl apply -f). Combine small, tested commands into scripts or Makefile targets to avoid manual drift. The short-term pain of learning single-command idioms (e.g., kubectl port-forward, kubectl rollout status) saves hours in debugging and incident response.
Example compact cheat sheet (copy into your dotfiles or repo readme):
- Git: git status, git add -p, git commit -m, git rebase -i, git push –force-with-lease
- Docker: docker build -t app:latest ., docker run -it –rm app:latest, docker images, docker system prune
- Kubectl: kubectl apply -f, kubectl get pods -o wide, kubectl logs -f, kubectl exec -it
- Terraform: terraform init, terraform plan -out=plan.tfplan, terraform apply plan.tfplan, terraform fmt
For an organized collection of commands and examples, see the practical repository on GitHub: Kubernetes manifests & Terraform module scaffold examples.
Cloud infrastructure skills: beyond the CLI
CLI fluency is necessary but not sufficient. You must understand cloud service models (IaaS, PaaS, FaaS), networking basics (VPC, subnets, routing), identity and access management (IAM roles/policies), and cost-aware resource sizing. These concepts determine how you design secure, scalable systems and which commands or IaC patterns to use.
Practice mapping components to managed cloud services: databases (RDS/Cloud SQL), caches (ElastiCache), and object storage (S3/GCS). Learn provisioning patterns: immutable infra for stateless services, stateful sets for databases, and blue/green or canary strategies for deployments. Each pattern simplifies incident response and rollback.
Automate account-wide guardrails: use organization policies, IaC-based IAM, and automated compliance checks. Store ephemeral secrets in vaults and use short-lived credentials where possible. These are skills that pair with commands to keep production healthy and auditable.
CI/CD pipelines: best practices and commands
CI/CD is where commands meet automation. Pipelines must be deterministic and fast. Key actions: checkout code, run lint/tests, build artifacts, run security scans, and deploy. Learn the YAML syntax or DSL for your CI tool (GitHub Actions, GitLab CI, Jenkins, CircleCI) and keep pipeline stages small and parallelizable.
Implement gating: require successful tests and security scans before merging. Use artifact registries (container registries, package registries) and immutable tags for deployable units. Practice local runs of pipeline steps with containerized runners so you can iterate quickly without waiting on the full pipeline.
Example CI snippets are included in the linked repo; search for pipeline templates and adapt them to your branching strategy. Use ephemeral environments for pull requests to validate integration before merge — this reduces regressions and speeds up time-to-merge.
Container orchestration: production patterns with Kubernetes
Kubernetes is powerful but opinionated. Production manifests should be composable and environment-aware. Prefer Deployments for stateless workloads, StatefulSets for services requiring stable identities, and DaemonSets for node-level agents. Use ConfigMaps and Secrets for configuration and RBAC for least privilege.
Key production concerns include resource requests/limits, probe configuration, and pod disruption budgets. Requests ensure the scheduler places pods correctly; limits prevent noisy neighbors from starving other workloads. Liveness and readiness probes let Kubernetes detect unhealthy pods and stop routing traffic while a pod is still performing startup tasks.
For templating and environment overlays prefer Helm, Kustomize, or plain YAML with a CI-driven templating step. Store canonical manifests in a GitOps repo and let an operator (Argo CD/Flux) reconcile live clusters with the declared state. Examples of effective Kubernetes manifests and patterns are available in the repository: Kubernetes manifests.
Infrastructure as Code (IaC): writing reusable Terraform modules
Terraform modules must be opinionated, well-documented, and parameterized. A robust module scaffold includes main.tf, variables.tf (with types and defaults), outputs.tf, providers.tf, and a README.md with examples. Add examples/ for common consumption patterns and unit tests with tools like terratest or terraform validate in CI.
Design modules for composition rather than monoliths. Each module should own a single responsibility (VPC, RDS instance, IAM roles). Use clear naming and versioning (semantic versioning) so consumers can upgrade safely. Avoid embedding environment-specific values in modules — pass them as inputs or use workspaces carefully.
Maintain state hygiene: use remote backends with locking (S3 backend + DynamoDB lock, GCS + state locking plugin) and secure access (least-privilege IAM). For collaborating teams, include a migration plan for state changes and document how to import existing resources if needed. A canonical scaffold and examples are included in the repo; refer to the Terraform module scaffold for a tested layout.
Monitoring, observability, and incident response
Monitoring is more than metrics: it’s traces, logs, and alerting. Implement a three-layer observability stack: metrics (Prometheus), traces (OpenTelemetry/Jaeger), and logs (ELK/Fluentd/Cloud Logging). Ensure service-level indicators (SLIs) and service-level objectives (SLOs) are defined to guide alerting thresholds and reduce noise.
Alerting must be actionable. Define playbooks for common alerts (e.g., CPU/OOM, latency spikes, database connection errors) and link them to runbooks that include command snippets to triage and mitigate. Automate common remediation when safe — for example, autoscaling policies or circuit breakers — but avoid over-automation that obscures root causes.
Practice incident drills and postmortems with blameless culture. Use runbooks and documented CLI commands to accelerate recovery: kubectl exec for quick fixes, terraform plan for infra drift checks, and cloud CLI for account-level diagnostics. A short list of incident triage commands and patterns is available in the linked command repo.
Putting it together: workflows and examples
Real-world workflows combine all the above: a developer opens a PR, CI runs tests and builds a container, the registry stores the artifact, a deployment job updates the cluster using a Helm chart or kubectl apply, monitoring validates health, and alerts trigger if an SLO is breached. Each touchpoint should be automated and auditable.
Practice by building a small end-to-end project: app code repository → CI pipeline → container registry → Terraform infra scaffold for cluster and networking → Kubernetes manifests or Helm chart → GitOps reconciliation. Iterate on each step to reduce mean time to deploy and mean time to recovery.
If you want an actionable collection of commonly used commands, sample manifests, and Terraform module scaffolds to get started, clone this hands-on repo and adapt it: DevOps commands & scaffolds.
- Git, Docker, kubectl, Helm/Kustomize
- Terraform (or CloudFormation/ARM), terratest (optional)
- Prometheus/Grafana, OpenTelemetry, ELK or cloud logging
- CI runner: GitHub Actions / GitLab / Jenkins
Best practices checklist
Small set of repeatable heuristics that reduce incidents and increase deploy speed:
1) Keep manifests declarative and templatized; 2) Define SLOs and use them to shape alerts; 3) Version and test Terraform modules; 4) Use immutable artifacts and promote immutability in pipelines. Each rule reduces cognitive load and makes automation reliable.
Finally, document everything: commands that matter during incidents, how to read metrics and traces, and the location of infra code. Documentation plus automated checks is what makes a platform team sleep better.
FAQ
-
What are the essential DevOps commands I should know?
Focus on a handful across key tools: Git (git status, add, commit, rebase), Docker (build, run, ps, logs), kubectl (apply, get, logs, exec, port-forward), Terraform (init, plan, apply, fmt), and your cloud CLI (aws/az/gcloud basics). Learn commands that enable safe inspection and rollback; script them into CI/CD steps to avoid manual errors.
-
How do I structure a Terraform module scaffold?
Include main.tf, variables.tf with types and validation, outputs.tf, providers.tf, README.md, examples/, and optionally modules/. Use semantic versioning, default values sparingly, and remote state with locking. Add CI checks (terraform fmt, validate, plan) and unit tests (terratest) to enforce quality.
-
How do I write Kubernetes manifests for production?
Use Deployments or StatefulSets, set CPU/memory requests and limits, configure liveness/readiness probes, use ConfigMaps and Secrets for configuration, and apply RBAC policies. Prefer small, composable manifests and templating (Helm/Kustomize). Run pre-deployment checks and use GitOps for reconciliation to avoid drift.