A Spark Application that Runs Inside Your VPC
Without network_configuration the application runs on AWS-managed networking: it cannot reach a private database, and its egress skips your routing, NAT and DNS firewall. Monitoring is absent by default too, so a failed job leaves no executor logs, and no maximum_capacity makes the account quota the only ceiling.
Verification
Static-verifiedPassed: validated and lint-clean (provider-schema-validated for AWS/Azure/GCP; Terraform-language lint elsewhere).
Conformance
- Static validation (fmt · validate · tflint)
- Security scan clean (Checkov)
- Plan tests (mocked: validation rules · outputs)
Provenance
- SHA-256 checksum
- Signature (pending)
Functional
- Live test pending (no cloud run yet)
Last verified 2026-09-13 · how we verify
Use it from the registry
terraform · opentofumodule "emr_serverless" {
source = "www.iac-bazaar.com/iac-bazaar/aws-emr-serverless/aws"
version = "1.0.0"
}Needs a registry token from /account/tokens. The module itself is free; the account is what identifies you. Full setup: registry docs.
Inputs & outputs
Create a free account to read this module's contract
The declared contract - every input name, type, default and description, plus every output - is shown to signed-in accounts, not to anonymous visitors.
A free account sees the contract of every module in the catalogue. There is no subscription and nothing to buy - the modules are free to download, and they run under Vizier.
Documentation
aws-emr-serverless
An EMR Serverless application that runs inside your VPC, keeps its logs, and has
a ceiling. Works with Terraform and OpenTofu (>= 1.6), AWS provider
>= 6.0, < 7.0.
Without network_configuration the application runs on AWS-managed
networking, and that means two things at once: it cannot reach anything
private - not a database in your VPC, not a private Kafka, not an interface
endpoint - and its outbound traffic does not pass through your routing, your NAT
or your DNS firewall. Jobs that process your data then run somewhere you do not
inspect. A precondition refuses it unless accept_aws_managed_network says
it was meant, and a second one refuses half a network configuration, since
subnets without security groups is not a partial answer but a broken one.
Monitoring is off by default. No monitoring_configuration means the driver
and executor logs of a failed job go nowhere - and a Spark failure is diagnosed
from executor logs, not from the job status. On here, with a KMS key, because
the logs of a job that reads your data contain parts of your data.
maximum_capacity absent means the account service quota is the limit. A job
with a bad join scales until it reaches that, and the first symptom is the bill.
An explicit ceiling costs nothing to set and is the only thing that bounds it.
Pre-initialised capacity is deliberately not configured. initial_capacity
keeps workers warm so the next job starts immediately, and bills for them
whether or not a job ever arrives. That is a real trade-off rather than a
default, so the module leaves it off and says so here instead of setting it
quietly.
Smaller things: ARM64 rather than AWS's X86_64, because Graviton is
materially cheaper for Spark on the same managed runtime image; auto_stop on,
with a precondition that refuses an always-warm application unless
accept_always_on states it; and the CloudWatch block's enabled is a plain
variable rather than a derived expression, so a configuration scanner can
actually read it.
Verification
Static validation runs tofu fmt, init, validate, tflint and checkov.
This module has not yet had a live test, so it is published as statically
validated with its live test pending and does not carry the live-tested mark.
Usage code & full reference need an account
The complete copy-paste usage, the full input/output reference, and operational notes are free with an account - shown here and bundled in the download. Sign in and this section fills in.
- Usage
Related modules
aws-keyspaces
point_in_time_recovery defaults to DISABLED and Keyspaces has no snapshots or automated backups, so off means a dropped table is simply gone. PITR on, a customer-managed key, and the two one-way doors - client-side timestamps and TTL - named rather than set quietly.
aws-neptune
Neptune has no user, no password and no GRANT. Authorization is IAM and it defaults to OFF, so anything that can reach port 8182 can read every edge and drop the lot. IAM auth on, storage encrypted, and the audit log driven from one variable because its two halves live in different resources and either alone logs nothing.
aws-dms
ssl_mode defaults to none in AWS, so a task reads your entire production database and writes it elsewhere unencrypted. This defaults to require, refuses none unless stated, and pushes the credential into Secrets Manager rather than state.
aws-athena
Query results are a copy of the data, written to S3. Without enforce_workgroup_configuration - the AWS default - a client sends its own location and encryption and every setting becomes a suggestion.
aws-aurora
Aurora PostgreSQL/MySQL cluster with instances, parameter groups, Serverless v2 scaling, and enhanced monitoring.
aws-documentdb
A cluster whose two dangerous AWS defaults are inverted: storage encryption is hard-coded on because it cannot be added later, and the master password is never an input - Secrets Manager generates it, so it never reaches the state file.