Infrastructure as Code: Managing Servers With Terraform and Pulumi 

Infrastructure as Code: Managing Servers With Terraform and Pulumi 

A platform team at a healthcare startup once spent an entire weekend trying to reproduce a production incident in staging, only to discover the two environments had quietly drifted apart over eighteen months of manual fixes. Someone had bumped a security group rule directly in the AWS console during a previous incident and never documented it. Someone else had resized an instance type by hand to fix a memory issue and forgotten to update the staging equivalent.

The team could no longer answer a basic question what does production really look like right now without logging into the console and clicking through every resource by hand.

That weekend became the case for infrastructure as code, and specifically for treating infrastructure changes with the same discipline as application code changes: reviewed, versioned, and applied through a tool rather than a person clicking buttons. 

An Outage Caused by a Manual Console Change 

Manually provisioned infrastructure accumulates undocumented state the same way an unmaintained codebase accumulates technical debt, except it’s often worse, because there’s rarely a diff to review, a commit history to trace, or a rollback path beyond someone’s memory of what they clicked. Every manual change to a running system is a small bet that nobody will need to know later exactly what changed and why. 

Infrastructure as code treats server configuration, networking, and cloud resources as version-controlled artifacts, written in a declarative or programmatic language, applied through a tool that computes and executes the necessary changes. The immediate benefits are the same ones code review brought to application development: changes are visible before they happen, reviewable by a second person, and reproducible exactly the same way every time, whether it’s the tenth environment or the hundredth.

  • Reproducibility: standing up an identical environment becomes a matter of running the same code, not remembering a sequence of manual steps. 
  • Auditability: every infrastructure change has an author, a timestamp, and a reason, tied to version control history. 
  • Disaster recovery: rebuilding infrastructure after a catastrophic failure becomes a matter of re-running known-good code rather than reverse-engineering what existed. 
  • Consistency across environments: staging and production can be defined from the same underlying code, with only parameterized differences. 

Declarative vs Imperative Provisioning 

Infrastructure as code tools split broadly into declarative and imperative approaches, and the distinction shapes how you think about changes over the system’s lifetime. Declarative tools ask you to describe the desired end state “there should be three web servers, a load balancer, and a database” and the tool figures out what actions are needed to get from the current state to that description. 

# Terraform (declarative HCL) 

resource "aws_instance" "web" { 

count = 3 

ami = "ami-0abcdef1234567890" 

instance_type = "t3.medium" 

tags = { 

Name = "web-server-${count.index}" 

} 

}

Imperative tools, by contrast, ask you to describe the steps to take “create a server, then install nginx, then configure it” closer to a traditional script. Ansible is the most widely used example, though it supports idempotent operations that make repeated runs safe, blurring the line somewhat with declarative tools. 

Most modern IaC practice leans declarative for provisioning (Terraform, Pulumi, CloudFormation) and imperative for configuration management within already-provisioned servers (Ansible, Chef), using each where it fits best rather than forcing one paradigm to cover the entire stack. 

Terraform State Management 

Terraform State Management

Terraform’s defining mechanic is its state file, a record of what resources it believes exist and their current configuration, which it uses to compute the difference between desired and actual state on every run. This state file is central to how Terraform works, and mismanaging it is one of the most common sources of real production incidents in teams new to the tool. 

By default, state lives in a local file, which works for a single engineer experimenting but breaks down immediately with more than one contributor, since two people running Terraform against their own local state files can produce conflicting, divergent views of the same infrastructure. Production setups store state remotely commonly in an S3 bucket with DynamoDB-based locking, or in Terraform Cloud so every contributor works against the same source of truth, and locking prevents two applies from running concurrently and corrupting that shared state. 

  • Remote state backend: a shared, durable location for the state file, accessible to every team member and CI pipeline. 
  • State locking: preventing concurrent applies from racing against each other and corrupting the state file. 
  • State drift: when actual infrastructure diverges from what the state file records, often from a manual change made outside Terraform. 
  • Workspace isolation: separate state files per environment (staging, production), avoiding accidental cross-environment changes. 

Drift detection periodically running `terraform plan` against production even without an intended change catches the exact problem the healthcare startup ran into: someone changing something by hand outside the tool’s awareness, silently invalidating the assumption that the code describes reality. 

State file corruption is a related, sharper-edged risk worth planning for explicitly. Because every apply reads and rewrites the state file, an interrupted apply a network failure mid-run, a CI job killed abruptly can occasionally leave the state file in a partially updated condition that doesn’t cleanly match either the old or new desired configuration.

Remote backends with locking reduce the chance of two applies colliding, but they don’t eliminate the risk of a single interrupted apply leaving inconsistent state behind. Regular state backups, combined with a documented recovery procedure using terraform import or the state backend’s own versioning (S3 bucket versioning is a common, low-effort safety net), turn what could be a multi-day incident into a contained, well-defined recovery task. 

Pulumi and Programming-Language Infrastructure 

Pulumi takes a different approach to declarative provisioning, letting engineers write infrastructure definitions in a general-purpose language TypeScript, Python, Go, or others rather than a domain-specific configuration language like Terraform’s HCL. The underlying model is similar: describe desired state, let the tool compute and apply the difference. 

import pulumi_aws as aws 

web_servers = [] 

for i in range(3): 

instance = aws.ec2.Instance(f"web-server-{i}", 

ami="ami-0abcdef1234567890", 

instance_type="t3.medium", 

tags={"Name": f"web-server-{i}"}) 

web_servers.append(instance)

The appeal is access to a full programming language’s tooling loops, conditionals, functions, existing package ecosystems, and IDE support like autocomplete and type checking rather than working within a configuration language’s more limited expressiveness. Teams that already have strong conventions around code review, testing, and modularity in their application codebase often find that familiarity extends naturally to infrastructure code written in the same language. 

  • Familiar language ecosystem: reusing existing linters, test frameworks, and package managers rather than learning HCL-specific tooling. 
  • Stronger abstraction capabilities: functions and classes can encapsulate infrastructure patterns more naturally than HCL modules. 
  • Type safety: catching certain classes of configuration mistakes at compile time or through IDE tooling, before a plan is even run. 
  • Steeper learning curve for infrastructure-only teams: engineers without programming backgrounds may find HCL’s narrower, purpose-built syntax easier to pick up initially. 

Modules, Reuse, and Team Workflows 

Neither Terraform nor Pulumi is meant to be used as one enormous file describing every resource in an organization both support modular composition, where a well-tested pattern (a standard VPC setup, a common database configuration) is packaged once and reused across many projects and teams. 

Terraform modules are directories of configuration that accept input variables and expose output values, callable from other configurations much like a function. Pulumi achieves similar reuse through ordinary language constructs classes, functions, and packages that can be published and imported the same way application code libraries are.

Both approaches let a platform team codify organizational standards (required tags, approved instance types, mandatory encryption settings) once, and have every team consuming the module inherit those standards automatically rather than reimplementing them. 

  • Golden path modules: pre-approved, well-tested infrastructure patterns that most teams should default to rather than writing resources from scratch. 
  • Version pinning: modules referenced by a specific version rather than a moving branch, so an update to the module doesn’t silently change every consumer’s infrastructure.
  • Module registries: internal or public catalogs (the Terraform Registry, a private Pulumi package feed) that make discovering and adopting existing patterns easier than reinventing them.
  • Clear ownership boundaries: modules typically owned by a platform or infrastructure team, with application teams consuming rather than modifying them directly. 

Drift Detection and Remediation 

Even with disciplined IaC practice, drift happens an emergency fix applied by hand during an incident, a setting changed through a cloud console because someone forgot the proper workflow under pressure, or a resource modified by an entirely separate automated process the IaC tool doesn’t know about. Left unaddressed, drift undermines the entire premise of infrastructure as code, since the code no longer accurately describes what’s really running. 

Regular drift detection scheduled `terraform plan` runs, or Pulumi’s equivalent preview command, checked against production without applying anything surfaces these discrepancies before they cause confusion during the next legitimate change. Some teams wire this into CI, running a nightly plan and alerting if it detects unexpected differences, treating drift the same way a testing pipeline treats a failing test: something to investigate immediately rather than something to notice weeks later when it’s already caused a problem. 

Remediation typically takes one of two paths: updating the code to match the manual change if it was a legitimate, intentional fix that should persist, or reverting the manual change to match the code if it wasn’t meant to be permanent. Either resolution is better than leaving the discrepancy unresolved, since an unresolved drift means the next apply might silently undo someone’s emergency fix, or worse, produce unexpected behavior nobody predicted. 

Secrets and Sensitive Values in IaC 

Secrets and Sensitive Values in IaC

Infrastructure code frequently needs access to sensitive values database passwords, API keys, TLS certificates and handling these carelessly is one of the more common security mistakes teams make when adopting IaC. Committing a plaintext secret into version control, even briefly, leaves it in the repository’s history essentially forever, since removing a file doesn’t remove it from prior commits without a deliberate history rewrite. 

The standard practice is keeping secrets out of the code entirely, instead referencing them from a dedicated secrets manager (AWS Secrets Manager, HashiCorp Vault, Google Secret Manager) at apply time, so the actual secret value never appears in the repository or, ideally, in the state file either. Terraform’s state file itself deserves particular caution here, since it can end up containing resolved secret values in plaintext depending on the resource type, making a properly access-controlled and encrypted remote state backend a security requirement, not just an operational convenience. 

  • External secrets managers: the source of truth for sensitive values, referenced rather than embedded in code. 
  • Encrypted state storage: state backends configured with encryption at rest, given the sensitive data that can end up there. 
  • Least-privilege apply credentials: the identity running `terraform apply` or `pulumi up` scoped to only the permissions it really needs. 
  • Secret rotation compatibility: designing infrastructure code so rotating a secret doesn’t require a manual, error-prone process outside the normal workflow. 

Testing and Validating Infrastructure Changes

Treating infrastructure like application code implies testing it like application code, though the tooling for this is less mature than for typical application testing. Static analysis tools like `tflint` or `checkov` catch common mistakes and policy violations an overly permissive security group, an unencrypted storage bucket before a change is even applied, functioning similarly to a linter in a conventional codebase. 

Policy-as-code frameworks like Open Policy Agent or Sentinel let organizations codify rules (“no public S3 buckets,” “all instances must have required tags”) and enforce them automatically as part of the apply pipeline, rejecting changes that violate policy before they ever reach production. Integration testing is harder, since spinning up real cloud resources to validate a change is slower and costlier than a typical unit test, but tools like Terratest allow provisioning real infrastructure in a disposable test account, verifying it behaves as expected, and tearing it down automatically. 

A pragmatic middle layer many teams settle on is a staged rollout for infrastructure changes, mirroring how application deployments often work. A change is first applied against a low-stakes sandbox account, then a staging environment that mirrors production as closely as budget allows, and only then against production itself, with each stage’s plan output reviewed before proceeding.

This doesn’t replace automated testing, but it catches a category of mistake automated checks often miss a change that’s syntactically valid and policy-compliant but still produces an outcome nobody really intended, like an instance type change that technically satisfies every rule yet quietly doubles the monthly bill for a service nobody meant to resize. 

Final Thoughts 

Infrastructure as code is less about any specific tool and more about a discipline: treating servers, networks, and cloud resources as artifacts with a known, reviewable, reproducible definition, rather than a state of affairs only a handful of engineers can describe from memory.

Terraform’s declarative HCL and Pulumi’s general-purpose language approach both get you there, and the choice between them often comes down to team preference and existing tooling investment more than a decisive technical advantage on either side.

What matters more than the tool is the surrounding practice remote state, drift detection, policy enforcement, and a review process that catches mistakes before they reach production. Teams that adopt that discipline consistently rarely spend a weekend, the way the healthcare startup did, trying to reverse-engineer what production really looks like.

Frequently Asked Questions 

1. Can Terraform and Pulumi be used together in the same organization? 

They can coexist, often by team or by project, though managing two tools adds cognitive overhead and duplicated tooling investment. Most organizations standardize on one to keep module reuse, training, and CI pipelines consistent across teams. 

2. How does infrastructure as code handle resources that already exist outside the tool? 

Both Terraform and Pulumi support importing existing resources into their state, bringing infrastructure that was originally created by hand under the tool’s management going forward. The import process typically requires manually writing the corresponding resource definition to match the real resource’s current configuration, which can be tedious for a large, long-neglected environment but is generally a one-time cost. 

3. What happens if the state file is lost? 

It can often be reconstructed through import commands that reconcile existing cloud resources back into a new state file, but this process is manual, error-prone, and best avoided entirely through reliable, backed-up remote state storage rather than relied upon as a recovery plan.

4. Does infrastructure as code eliminate the need for a change approval process? 

No, it changes the shape of that process rather than removing it. Code review of a pull request replaces (or supplements) manual sign-off meetings, and the `plan` output before an `apply` gives reviewers a precise, machine-generated preview of exactly what will change. 

5. Is Ansible a replacement for Terraform? 

Generally no they’re complementary. Terraform (or Pulumi) typically provisions the underlying infrastructure (servers, networks, databases), while Ansible configures software and settings on top of already-provisioned servers, and many pipelines run both in sequence. 

6. Should every infrastructure change go through a pull request? 

For anything beyond personal experimentation, yes. Even small changes benefit from the review, plan preview, and audit trail a pull request workflow provides, and establishing that discipline early avoids the gradual reversion to manual changes that caused the drift problem in the first place. 

Similar Posts