Operations & Monitoring
Monitoring, backup, disaster recovery, and operational procedures for PeopleHub.
Monitoring & Logging
CloudWatch Logs
Lambda Logs: All Lambda invocations logged to CloudWatch Log Format: Structured JSON Retention: 90 days Log Groups: Separate for each Lambda function
Log Groups:
/aws/lambda/peoplehub-api-dev-api/aws/lambda/candidate-api-dev-api/aws/lambda/integration-api-dev-api- etc.
Queryable via CloudWatch Insights:
- Error analysis
- Performance debugging
- User activity tracking
CloudWatch Metrics
Lambda Metrics:
- Invocations, duration, errors, throttles
- Concurrent executions
- Cold start frequency
RDS Metrics:
- CPU utilization
- Database connections
- Read/write IOPS
- Disk space
API Gateway Metrics:
- Request count
- Latency (p50, p95, p99)
- 4xx/5xx errors
Planned Monitoring
X-Ray Tracing: End-to-end request tracing Custom Dashboards: Real-time system health view CloudWatch Alarms: Proactive alerts for:
- High error rates
- Elevated latency
- RDS connection exhaustion
- Lambda throttling
Backup & Recovery
RDS Database Backups
Automated Daily Snapshots:
- Retention: 30 days
- Encrypted with AWS KMS
- Stored in S3
Point-in-Time Recovery (PITR):
- Transaction log backups every 5 minutes
- Restore to any second in last 35 days
- RTO (Recovery Time Objective): <15 minutes
- RPO (Recovery Point Objective): <5 minutes
Manual Snapshots:
- Pre-deployment backups
- Before major data migrations
- Indefinite retention until manually deleted
S3 Document Backups
Versioning: Enabled on document buckets (restore previous versions) Cross-Region Replication: Planned for production Lifecycle Policies: Archive old versions to Glacier after 90 days
Lambda Code Backups
Version Control: All code in GitHub Lambda Versions: Each deployment creates new version Rollback: Update Lambda alias to previous version (instant)
Disaster Recovery
Recovery Metrics
RTO (Recovery Time Objective): <15 minutes RPO (Recovery Point Objective): <5 minutes
DR Scenarios
1. Database Failure (AZ Outage):
- Multi-AZ automatic failover (<30 seconds)
- No manual intervention required
- No data loss
2. Complete Region Outage:
- Restore RDS from snapshot in different region
- Deploy Lambda functions to new region
- Update DNS (Route 53)
- RTO: 30-60 minutes
3. Data Corruption/Deletion:
- Point-in-time recovery to before incident
- Validate data in test environment
- Cutover to restored database
- RTO: 15-30 minutes
4. Accidental Configuration Change:
- Infrastructure as code (revert git commit)
- Redeploy via Serverless Framework
- RTO: 10-20 minutes
Multi-Region DR (Planned for Global Offices)
Setup:
- Primary region: ap-south-1 (Mumbai)
- DR region: us-east-1 or eu-west-1
- Read replicas in DR region
- Route 53 failover routing
Failover: Automatic or manual promotion of read replica
Audit Trails
What's Audited
User Actions:
- Login/logout
- Data modifications (CRUD operations)
- Permission changes
- Configuration updates
System Events:
- Deployments
- Security events
- Integration failures
Audit Log Storage
Database: Audit tables with user, timestamp, action, old/new values CloudWatch: Application-level logs CloudTrail: AWS API call logs (planned)
Retention: 365 days for compliance
Compliance Reporting
Reports Available:
- User access logs
- Data modification history
- Security incident log
- System availability report
Incident Response
Incident Types
1. Service Outage:
- Check CloudWatch metrics
- Review recent deployments
- Rollback if needed
- Scale up resources if capacity issue
2. Performance Degradation:
- Check RDS connections and CPU
- Review slow queries (RDS Performance Insights)
- Optimize queries or scale up RDS instance
3. Security Incident:
- Isolate affected systems
- Analyze CloudTrail/CloudWatch logs
- Rotate credentials
- Notify stakeholders
4. Data Loss/Corruption:
- Immediately stop writes (if ongoing corruption)
- Restore from backup (PITR or snapshot)
- Validate data integrity
- Investigate root cause
Escalation Path
- Development team (first responders)
- Tech lead / DevOps
- CTO / Client technical team
Communication Plan
- Internal: Slack/email alerts
- Client: Email notification for production incidents
- Status page (future): Public status updates
Operational Runbooks
Common Tasks
1. Deploy to Production:
- Merge to
mainbranch - GitHub Actions triggers deployment
- Monitor CloudWatch for errors
- Verify health endpoint
- Rollback if issues
2. Scale Up RDS:
- Identify instance size needed
- Schedule maintenance window (few minutes downtime)
- Apply instance class change
- Monitor performance post-change
3. Rotate Secrets:
- Generate new secret in Secrets Manager
- Update all services (zero-downtime deployment)
- Validate new secret works
- Delete old secret after 24 hours
4. Investigate High Error Rate:
- Check CloudWatch metrics for spike
- Query logs for error details
- Identify affected endpoint
- Fix issue or rollback deployment
Integrations & External Systems
Current Integrations
Digio (E-Signature Platform):
- Digital signatures for onboarding policies
- Integration: Webhook + API
- Authentication: Client ID + Client Secret
- Webhook:
POST /webhooks/digio - Status: Production-ready
Planned Integrations
TalentRecruit (ATS): Candidate data sync (daily scheduled) Payroll Systems: Employee and salary data sync (weekly/monthly) Workforce Management: Project time tracking sync Background Verification: Automated background checks
Integration Architecture
Centralized Service: Integration API handles all external integrations Patterns: Webhook push, scheduled polling, real-time API Security: API keys in Secrets Manager, webhook signature validation
Planned Operational Enhancements
Automated Alerts: CloudWatch alarms for critical metrics On-Call Rotation: PagerDuty integration (for production) Automated Remediation: Lambda functions to auto-fix common issues Capacity Planning: Predictive scaling based on usage patterns Performance Optimization: Regular query optimization reviews