Skip to content

Operations & Monitoring ​

Monitoring, backup, disaster recovery, and operational procedures for PeopleHub.


Monitoring & Logging ​

CloudWatch Logs ​

Lambda Logs: All Lambda invocations logged to CloudWatch Log Format: Structured JSON Retention: 90 days Log Groups: Separate for each Lambda function

Log Groups:

  • /aws/lambda/peoplehub-api-dev-api
  • /aws/lambda/candidate-api-dev-api
  • /aws/lambda/integration-api-dev-api
  • etc.

Queryable via CloudWatch Insights:

  • Error analysis
  • Performance debugging
  • User activity tracking

CloudWatch Metrics ​

Lambda Metrics:

  • Invocations, duration, errors, throttles
  • Concurrent executions
  • Cold start frequency

RDS Metrics:

  • CPU utilization
  • Database connections
  • Read/write IOPS
  • Disk space

API Gateway Metrics:

  • Request count
  • Latency (p50, p95, p99)
  • 4xx/5xx errors

Planned Monitoring ​

X-Ray Tracing: End-to-end request tracing Custom Dashboards: Real-time system health view CloudWatch Alarms: Proactive alerts for:

  • High error rates
  • Elevated latency
  • RDS connection exhaustion
  • Lambda throttling

Backup & Recovery ​

RDS Database Backups ​

Automated Daily Snapshots:

  • Retention: 30 days
  • Encrypted with AWS KMS
  • Stored in S3

Point-in-Time Recovery (PITR):

  • Transaction log backups every 5 minutes
  • Restore to any second in last 35 days
  • RTO (Recovery Time Objective): <15 minutes
  • RPO (Recovery Point Objective): <5 minutes

Manual Snapshots:

  • Pre-deployment backups
  • Before major data migrations
  • Indefinite retention until manually deleted

S3 Document Backups ​

Versioning: Enabled on document buckets (restore previous versions) Cross-Region Replication: Planned for production Lifecycle Policies: Archive old versions to Glacier after 90 days

Lambda Code Backups ​

Version Control: All code in GitHub Lambda Versions: Each deployment creates new version Rollback: Update Lambda alias to previous version (instant)


Disaster Recovery ​

Recovery Metrics ​

RTO (Recovery Time Objective): <15 minutes RPO (Recovery Point Objective): <5 minutes

DR Scenarios ​

1. Database Failure (AZ Outage):

  • Multi-AZ automatic failover (<30 seconds)
  • No manual intervention required
  • No data loss

2. Complete Region Outage:

  • Restore RDS from snapshot in different region
  • Deploy Lambda functions to new region
  • Update DNS (Route 53)
  • RTO: 30-60 minutes

3. Data Corruption/Deletion:

  • Point-in-time recovery to before incident
  • Validate data in test environment
  • Cutover to restored database
  • RTO: 15-30 minutes

4. Accidental Configuration Change:

  • Infrastructure as code (revert git commit)
  • Redeploy via Serverless Framework
  • RTO: 10-20 minutes

Multi-Region DR (Planned for Global Offices) ​

Setup:

  • Primary region: ap-south-1 (Mumbai)
  • DR region: us-east-1 or eu-west-1
  • Read replicas in DR region
  • Route 53 failover routing

Failover: Automatic or manual promotion of read replica


Audit Trails ​

What's Audited ​

User Actions:

  • Login/logout
  • Data modifications (CRUD operations)
  • Permission changes
  • Configuration updates

System Events:

  • Deployments
  • Security events
  • Integration failures

Audit Log Storage ​

Database: Audit tables with user, timestamp, action, old/new values CloudWatch: Application-level logs CloudTrail: AWS API call logs (planned)

Retention: 365 days for compliance

Compliance Reporting ​

Reports Available:

  • User access logs
  • Data modification history
  • Security incident log
  • System availability report

Incident Response ​

Incident Types ​

1. Service Outage:

  • Check CloudWatch metrics
  • Review recent deployments
  • Rollback if needed
  • Scale up resources if capacity issue

2. Performance Degradation:

  • Check RDS connections and CPU
  • Review slow queries (RDS Performance Insights)
  • Optimize queries or scale up RDS instance

3. Security Incident:

  • Isolate affected systems
  • Analyze CloudTrail/CloudWatch logs
  • Rotate credentials
  • Notify stakeholders

4. Data Loss/Corruption:

  • Immediately stop writes (if ongoing corruption)
  • Restore from backup (PITR or snapshot)
  • Validate data integrity
  • Investigate root cause

Escalation Path ​

  1. Development team (first responders)
  2. Tech lead / DevOps
  3. CTO / Client technical team

Communication Plan ​

  • Internal: Slack/email alerts
  • Client: Email notification for production incidents
  • Status page (future): Public status updates

Operational Runbooks ​

Common Tasks ​

1. Deploy to Production:

  • Merge to main branch
  • GitHub Actions triggers deployment
  • Monitor CloudWatch for errors
  • Verify health endpoint
  • Rollback if issues

2. Scale Up RDS:

  • Identify instance size needed
  • Schedule maintenance window (few minutes downtime)
  • Apply instance class change
  • Monitor performance post-change

3. Rotate Secrets:

  • Generate new secret in Secrets Manager
  • Update all services (zero-downtime deployment)
  • Validate new secret works
  • Delete old secret after 24 hours

4. Investigate High Error Rate:

  • Check CloudWatch metrics for spike
  • Query logs for error details
  • Identify affected endpoint
  • Fix issue or rollback deployment

Integrations & External Systems ​

Current Integrations ​

Digio (E-Signature Platform):

  • Digital signatures for onboarding policies
  • Integration: Webhook + API
  • Authentication: Client ID + Client Secret
  • Webhook: POST /webhooks/digio
  • Status: Production-ready

Planned Integrations ​

TalentRecruit (ATS): Candidate data sync (daily scheduled) Payroll Systems: Employee and salary data sync (weekly/monthly) Workforce Management: Project time tracking sync Background Verification: Automated background checks

Integration Architecture ​

Centralized Service: Integration API handles all external integrations Patterns: Webhook push, scheduled polling, real-time API Security: API keys in Secrets Manager, webhook signature validation


Planned Operational Enhancements ​

Automated Alerts: CloudWatch alarms for critical metrics On-Call Rotation: PagerDuty integration (for production) Automated Remediation: Lambda functions to auto-fix common issues Capacity Planning: Predictive scaling based on usage patterns Performance Optimization: Regular query optimization reviews