AWS
Configure CloudWatch Logs and Alarms for Spring Boot Applications
By Utility Zone · 2026-10-03T15:52:27.951995
AWS article series — Security, Monitoring & Reliability
Audience: Java/Spring Boot developers, tech leads, and DevOps engineers
Goal: Centralize Spring Boot logs in Amazon CloudWatch Logs and create actionable CloudWatch alarms for application and infrastructure signals.
Architecture note: This guide uses an EC2-based Spring Boot deployment as the concrete example. The same CloudWatch concepts apply to ECS/Fargate and EKS, but log collection configuration differs by compute platform.
1. What you will build
By the end, you will have:
- Spring Boot application logs written to standard output or a file.
- A CloudWatch agent collecting logs from an EC2 instance.
- A dedicated CloudWatch log group with retention configured.
- Metric filters that turn selected log patterns into CloudWatch metrics.
- CloudWatch alarms for application errors and EC2 CPU utilization.
- An SNS notification topic for alarm notifications.
- A short verification and troubleshooting checklist.
Architecture
The application emits logs. The CloudWatch agent forwards them to a log group. Metric filters can publish metrics from matching log events, while native AWS metrics such as EC2 CPU utilization can be alarmed directly. CloudWatch alarms can notify an SNS topic, which can deliver email or integrate with other notification systems.
2. Prerequisites
- An AWS account and an EC2 instance running a supported Linux distribution.
- A Spring Boot application deployed on the instance.
- AWS CLI access for an administrator or deployment engineer.
- An EC2 instance profile / IAM role for the CloudWatch agent.
- Permission to create CloudWatch log groups, metric filters, alarms, and SNS topics.
- A tested email address if using SNS email notifications.
Use separate least-privilege identities for deployment automation and the EC2 workload. Avoid storing long-lived AWS access keys in application configuration or on the instance.
3. Decide how your application will write logs
For containerized deployments, writing logs to standard output is usually the simplest approach because the container platform can collect them. For this EC2 example, a file-based log is shown so that the CloudWatch agent has a concrete path to tail.
Spring Boot uses Logback by default when using the standard Spring Boot logging starter. A simple src/main/resources/logback-spring.xml can configure a rolling file appender. Adjust paths and retention to your operational requirements.
<?xml version="1.0" encoding="UTF-8"?>
<configuration>
<property name="LOG_DIR" value="${LOG_DIR:-/var/log/myapp}"/>
<appender name="FILE"
class="ch.qos.logback.core.rolling.RollingFileAppender">
<file>${LOG_DIR}/application.log</file>
<rollingPolicy class="ch.qos.logback.core.rolling.TimeBasedRollingPolicy">
<fileNamePattern>${LOG_DIR}/application.%d{yyyy-MM-dd}.log.gz</fileNamePattern>
<maxHistory>14</maxHistory>
</rollingPolicy>
<encoder>
<pattern>%d{yyyy-MM-dd'T'HH:mm:ss.SSSXXX} %-5level [%thread] %logger{36} - %msg%n</pattern>
</encoder>
</appender>
<root level="INFO">
<appender-ref ref="FILE"/>
</root>
</configuration>
Create the directory and ensure the OS account running the application can write to it:
sudo mkdir -p /var/log/myapp
sudo chown <app-user>:<app-group> /var/log/myapp
sudo chmod 750 /var/log/myapp
Replace <app-user> and <app-group> with the actual service account and group. Do not log passwords, access tokens, session identifiers, payment data, or unnecessary personal information.
If your app already writes structured JSON logs or uses a managed logback encoder, retain that configuration rather than replacing it blindly. Structured logs make filtering and analysis easier.
4. Create the CloudWatch log group
Choose a consistent naming convention, for example:
- Log group:
/myapp/prod/spring-boot - Log stream:
i-INSTANCE_ID/application
Create the log group and set a retention period. Retention should follow your organization's operational, privacy, and regulatory requirements.
aws logs create-log-group \
--log-group-name "/myapp/prod/spring-boot" \
--region ap-south-1
aws logs put-retention-policy \
--log-group-name "/myapp/prod/spring-boot" \
--retention-in-days 30 \
--region ap-south-1
The example uses Mumbai (ap-south-1); change the region to the one where your workload runs. The commands can return an error if the log group already exists, so check before creating it in repeatable automation.
5. Give the EC2 instance permission to publish logs
Attach an IAM role to the EC2 instance through an instance profile. For a quick initial setup, AWS provides the managed policy CloudWatchAgentServerPolicy. Review its permissions against your organization's security requirements before using it in production.
For a tighter production policy, scope permissions to the required log group and streams where practical. The agent generally needs permissions to describe log groups/streams and create streams and put log events. Exact permissions can vary with the agent configuration and whether it creates log groups.
Important: Do not put access keys in application.properties, environment files, user data, or the CloudWatch agent configuration. Use the instance role and temporary credentials.
6. Install the CloudWatch agent
On Amazon Linux 2023, install the package with DNF:
sudo dnf install -y amazon-cloudwatch-agent
On Amazon Linux 2, use the package manager available for that operating system image. For Ubuntu or other distributions, follow the AWS CloudWatch agent installation instructions for that OS.
Create a configuration file at:
/opt/aws/amazon-cloudwatch-agent/etc/amazon-cloudwatch-agent.json
Example configuration:
{
"agent": {
"run_as_user": "root",
"metrics_collection_interval": 60
},
"logs": {
"logs_collected": {
"files": {
"collect_list": [
{
"file_path": "/var/log/myapp/application.log",
"log_group_name": "/myapp/prod/spring-boot",
"log_stream_name": "{instance_id}/application",
"timezone": "UTC"
}
]
}
}
}
}
The root setting is shown as a simple example to avoid file-read permission issues; for production, prefer a dedicated agent user and grant only the access it needs. Confirm that the agent version supports the configuration fields you use.
Start the agent using the control script:
sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a fetch-config \
-m ec2 \
-c file:/opt/aws/amazon-cloudwatch-agent/etc/amazon-cloudwatch-agent.json \
-s
Check status and logs:
sudo /opt/aws/amazon-cloudwatch-agent/bin/amazon-cloudwatch-agent-ctl \
-a status
sudo tail -n 100 /opt/aws/amazon-cloudwatch-agent/logs/amazon-cloudwatch-agent.log
Generate a harmless test log from the application and confirm it appears in CloudWatch → Logs → Log groups → /myapp/prod/spring-boot.
7. Make error logs measurable with a metric filter
A CloudWatch Logs metric filter can count log events matching a pattern and publish a custom metric. The pattern must match the format your application actually emits.
For the sample text layout above, a filter pattern that matches ERROR can be:
aws logs put-metric-filter \
--log-group-name "/myapp/prod/spring-boot" \
--filter-name "SpringBootErrorCount" \
--filter-pattern '"ERROR"' \
--metric-transformations \
metricName=ApplicationErrorCount,metricNamespace=MyApp/Production,metricValue=1,defaultValue=0 \
--region ap-south-1
This metric counts matching log events, not necessarily unique incidents or failed user requests. A single exception may produce several log events depending on the logging format. Validate the filter with representative log samples before using it for paging.
If you use JSON logs, consider matching a structured field such as level = "ERROR" rather than searching arbitrary text. Test the filter pattern in the CloudWatch console with real, sanitized examples.
8. Create an alarm for application errors
Example: alert if at least five matching error log events are recorded within a five-minute period. This is illustrative; tune the threshold based on normal traffic, log volume, and incident response requirements.
aws cloudwatch put-metric-alarm \
--alarm-name "MyApp-ApplicationErrors" \
--alarm-description "Application error log count reached the configured threshold" \
--namespace "MyApp/Production" \
--metric-name "ApplicationErrorCount" \
--statistic Sum \
--period 300 \
--evaluation-periods 1 \
--datapoints-to-alarm 1 \
--threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--region ap-south-1
The command creates an alarm but does not attach a notification action yet. Add the SNS topic ARN after creating the topic, or include --alarm-actions <SNS_TOPIC_ARN> in the command.
Choose missing-data behavior deliberately. notBreaching can be appropriate for a count metric when no matching events means zero errors, but it is not suitable for every metric. A missing metric can also mean the agent or application has stopped reporting.
9. Create an SNS topic and subscribe an email address
aws sns create-topic \
--name "myapp-production-alerts" \
--region ap-south-1
Copy the returned topic ARN, then subscribe your operational email address:
aws sns subscribe \
--topic-arn "<SNS_TOPIC_ARN>" \
--protocol email \
--notification-endpoint "ops@example.com" \
--region ap-south-1
Replace the example email with the appropriate distribution list. The recipient must confirm the subscription from the email before notifications are delivered.
Attach the topic to the alarm:
aws cloudwatch put-metric-alarm \
--alarm-name "MyApp-ApplicationErrors" \
--alarm-description "Application error log count reached the configured threshold" \
--namespace "MyApp/Production" \
--metric-name "ApplicationErrorCount" \
--statistic Sum \
--period 300 \
--evaluation-periods 1 \
--datapoints-to-alarm 1 \
--threshold 5 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data notBreaching \
--alarm-actions "<SNS_TOPIC_ARN>" \
--region ap-south-1
Alarm actions run when the alarm transitions into the relevant state. Test the notification path in a controlled environment and document who owns the alert.
10. Add an infrastructure alarm for EC2 CPU
CloudWatch publishes EC2 CPU utilization metrics automatically for supported EC2 instances. This example alerts when average CPU is at least 80% for three consecutive five-minute periods.
aws cloudwatch put-metric-alarm \
--alarm-name "MyApp-EC2-HighCPU" \
--alarm-description "EC2 CPU utilization has remained high" \
--namespace "AWS/EC2" \
--metric-name "CPUUtilization" \
--dimensions Name=InstanceId,Value=<INSTANCE_ID> \
--statistic Average \
--period 300 \
--evaluation-periods 3 \
--datapoints-to-alarm 3 \
--threshold 80 \
--comparison-operator GreaterThanOrEqualToThreshold \
--treat-missing-data missing \
--alarm-actions "<SNS_TOPIC_ARN>" \
--region ap-south-1
Replace <INSTANCE_ID> and <SNS_TOPIC_ARN>. CPU is only one signal; a low CPU value does not prove the application is healthy. Consider application-level health checks, request latency, 5xx responses, JVM heap/GC metrics, disk space, and load balancer target health where relevant.
11. Optional: add an application health signal
For a Spring Boot application, Spring Boot Actuator can expose health information. Secure management endpoints and expose only what is required. A health endpoint returning success is not the same as monitoring request latency or business-level failures.
A common monitoring design includes:
- Availability: load balancer target health or an external synthetic check.
- Errors: HTTP 5xx count and application error metrics.
- Latency: p95/p99 request latency.
- Saturation: CPU, memory, JVM heap, thread pools, connection pools, and disk.
- Dependencies: database connection failures, timeouts, and downstream API errors.
If you are on ECS/Fargate, prefer the platform's awslogs log driver or an AWS-supported log router rather than installing the EC2 CloudWatch agent into a task without a specific reason. For EKS, choose a cluster logging architecture appropriate to the cluster and workload.
12. Verify end to end
- Confirm the application writes the expected log file.
- Confirm the CloudWatch agent status is
running. - Confirm the log group receives new events.
- Search for a known test message in CloudWatch Logs Insights.
- Confirm the metric filter is matching events and the custom metric appears.
- Temporarily lower the alarm threshold in a non-production environment or emit controlled test events.
- Confirm the alarm changes state and the SNS email arrives after subscription confirmation.
- Restore production thresholds and record the test result.
A basic Logs Insights query for text logs:
fields @timestamp, @message
| filter @message like /ERROR/
| sort @timestamp desc
| limit 100
For JSON logs, query parsed fields directly where possible. Use sanitized test data and avoid putting sensitive values in log messages.
13. Troubleshooting
| Symptom | Checks |
|---|---|
| No log stream appears | Check agent status and agent logs; verify file path, permissions, IAM role, region, and network access to CloudWatch endpoints. |
| Log stream exists but events are stale | Check whether the application is writing to the configured file; verify rotation behavior and the configured file path. |
| Metric filter shows no data | Confirm the pattern matches actual log events; test against sanitized samples; check namespace and region. |
Alarm stays in INSUFFICIENT_DATA | Confirm the metric is being published and review period, dimensions, evaluation settings, and missing-data behavior. |
| Alarm changes state but email is not received | Confirm the SNS subscription is confirmed, the topic ARN is correct, and the alarm has the topic in its actions. |
| Duplicate or noisy alerts | Review log-event semantics, thresholds, evaluation periods, and whether repeated stack-trace lines are counted separately. |
| Unexpected CloudWatch costs | Review ingestion volume, retention, metric cardinality, custom metrics, and dashboard/alarm usage. |
14. Production checklist
- Log group names distinguish application, environment, and service.
- Retention is explicitly configured and aligned with policy.
- IAM uses workload roles and least privilege.
- Secrets and sensitive personal data are excluded or redacted from logs.
- Log rotation and agent behavior have been tested.
- Metric filters match the actual logging format.
- Alarm thresholds are based on service behavior and operational objectives.
- Alarm actions notify an owned and monitored channel.
- Runbooks explain what to inspect when each alarm fires.
- Logs, metrics, and alarms are managed as code where practical.
- The team has tested notification delivery and recovery procedures.
15. Key takeaways
Centralized logs make it easier to investigate production behavior, while metric filters and CloudWatch alarms help convert selected signals into operational alerts. Reliable monitoring requires more than an alarm threshold: validate collection, metric semantics, notification delivery, retention, access controls, and the runbook that responders will follow.
Start with a small set of actionable alarms, observe their behavior, and refine them using production evidence. Avoid creating alerts that no team owns or that cannot lead to a clear action.
References
- Amazon CloudWatch Logs
- Install the CloudWatch agent
- CloudWatch agent configuration file
- Create metric filters for log groups
- Create CloudWatch alarms
- Amazon SNS email subscriptions
- Spring Boot logging
- Spring Boot Actuator
Educational example: validate commands, IAM permissions, operating-system package names, alarm thresholds, and logging configuration in a non-production environment before applying them to a production service.