Skip to content

Scalability Model ​

Scalability Overview ​

PeopleHub is designed to scale from 100 employees to 10,000+ without architectural changes, leveraging AWS serverless auto-scaling capabilities.

Component Scalability ​

Frontend (S3 + CloudFront) ​

Current Capacity: Unlimited Scaling Mechanism: Automatic (AWS-managed)

CloudFront edge locations automatically serve static assets globally:

  • 200+ edge locations worldwide
  • Automatic cache scaling
  • No configuration needed

Performance at Scale:

  • 100 users: <50ms latency
  • 10,000 users: <50ms latency (same)
  • No difference in frontend performance regardless of user count

API Gateway ​

Current Capacity: 10,000 requests/second (regional default) Scaling Mechanism: Automatic

  • Burst capacity: 5,000 requests
  • Sustained capacity: 10,000 requests/second
  • Can be increased via AWS support request (no limit)

No action required for scaling up to 10,000 concurrent users.

Lambda Functions ​

Current Capacity: 1,000 concurrent executions (default) Scaling Mechanism: Automatic, sub-second

How Lambda Scales:

  1. First request creates new Lambda instance
  2. Request processed (warm execution: 10-50ms)
  3. Subsequent requests reuse warm instance
  4. If all instances busy, AWS creates new instances automatically
  5. Scales from 0 to 1,000 concurrent executions in seconds

Scaling Example:

  • Normal: 50 concurrent requests → 50 Lambda instances
  • Peak: 500 concurrent requests → 500 Lambda instances (auto-scaled)
  • Night: 0 requests → 0 Lambda instances (scaled to zero, no cost)

Limits:

  • Default: 1,000 concurrent executions per region
  • Can be increased to 10,000+ via AWS support
  • No application changes needed

RDS PostgreSQL ​

Current Configuration: Single instance, Multi-AZ Scaling Mechanism: Vertical scaling (manual), read replicas (future)

Vertical Scaling:

  • Current: t3.medium (2 vCPU, 4 GB RAM) - sufficient for 1,000 employees
  • Scale up: t3.large, t3.xlarge, or m5/r5 instances
  • Scaling requires brief downtime (few minutes) for instance change

Connection Pool:

  • Each Lambda maintains 10 connections
  • RDS supports 100-1000+ connections based on instance size
  • RDS Proxy can be added for better connection management (if needed)

Read Replicas (future):

  • Create read-only replicas for reporting
  • Offload read traffic from primary database
  • Up to 15 read replicas supported

S3 Document Storage ​

Current Capacity: Unlimited Scaling Mechanism: Automatic

S3 automatically scales to handle:

  • Unlimited files
  • Unlimited storage
  • Thousands of requests per second

No configuration needed for scaling.

Scaling Scenarios ​

Scenario 1: Company Growth (100 → 1,000 Employees) ​

Frontend: No changes needed Lambda: Auto-scales to handle increased API traffic RDS: May need to scale up instance size S3: No changes needed

Action Required: Monitor RDS performance, scale up if needed

Scenario 2: High-Volume Onboarding (50+ Candidates/Month) ​

Candidate API Lambda: Auto-scales to handle concurrent onboarding S3: Auto-scales for document uploads Notifications: Auto-scales for email sends

Action Required: None, automatic scaling handles this

Scenario 3: Global Expansion (7 Offices Worldwide) ​

CloudFront: Already global, no changes RDS: Add read replicas in other regions Lambda: Deploy to multiple regions (optional for reduced latency)

Action Required:

  • Setup cross-region RDS read replicas
  • Configure Route 53 geolocation routing (optional)

Scenario 4: Peak Usage (All Employees Access Simultaneously) ​

Example: Company announcement at 9 AM, 5,000 employees log in

Traditional System:

  • Servers overload
  • 503 errors
  • 5-10 minutes to auto-scale

PeopleHub Serverless:

  • Lambda scales to 5,000 concurrent executions in <2 seconds
  • CloudFront serves frontend from cache (no backend load)
  • All users served successfully
  • No manual intervention

Performance Targets by Scale ​

UsersConcurrent RequestsLambda InstancesAPI LatencyRDS Instance
1001010<100mst3.small
5005050<150mst3.medium
1,000100100<200mst3.large
5,000500500<200msm5.xlarge
10,0001,0001,000<250msm5.2xlarge + read replicas

Cost Scaling ​

Cost Comparison at Different Scales ​

At 100 Employees (Low Traffic):

  • Traditional servers: $250/month (always running)
  • PeopleHub serverless: $60/month (pay-per-use)
  • Savings: 76%

At 1,000 Employees (Moderate Traffic):

  • Traditional servers: $600/month
  • PeopleHub serverless: $180/month
  • Savings: 70%

At 5,000 Employees (High Traffic):

  • Traditional servers: $2,000/month
  • PeopleHub serverless: $600/month
  • Savings: 70%

Key Insight: Serverless becomes MORE cost-effective at scale due to efficient resource utilization.

Scaling Limits & Mitigation ​

Lambda Concurrency Limit ​

Default: 1,000 concurrent executions Mitigation: Request increase via AWS support (can scale to 10,000+) When Needed: At 5,000+ concurrent users

RDS Connection Pool ​

Limit: Based on instance size (e.g., 500 connections for m5.large) Mitigation: Add RDS Proxy for connection pooling When Needed: At 2,000+ concurrent database operations

API Gateway Throttling ​

Default: 10,000 requests/second Mitigation: Request quota increase via AWS support When Needed: At 10,000+ sustained requests/second

Monitoring for Scalability ​

Key Metrics to Watch:

  • Lambda concurrent executions (alert at 80% of limit)
  • RDS CPU and connections (alert at 70% utilization)
  • API Gateway 4xx/5xx errors (alert on spike)
  • Application response times (p95 latency)

CloudWatch Alarms (planned):

  • Lambda throttling errors
  • RDS connection exhaustion
  • API Gateway throttling
  • High error rates

Auto-Scaling vs Manual Scaling ​

ComponentScaling TypeConfiguration Needed
CloudFrontAutoNone
API GatewayAutoNone
LambdaAutoNone (monitor limits)
RDSManualInstance size change
S3AutoNone

Only RDS requires manual intervention for vertical scaling.

Future Scalability Enhancements ​

Multi-Region Deployment ​

  • Deploy Lambda functions in multiple AWS regions
  • Read replicas in each region
  • Route 53 for geolocation-based routing
  • Reduces latency for global users

Aurora Serverless v3 ​

  • Reevaluate Aurora Serverless (better than v2)
  • Automatic database scaling (ACUs)
  • No manual instance sizing

ElastiCache (Redis) ​

  • Add caching layer for frequently accessed data
  • Reduce database load
  • Improve response times

SQS for Async Processing ​

  • Queue long-running operations
  • Process asynchronously via separate Lambda
  • Improve API response times

Scalability Best Practices ​

  1. Stateless Design: No session state in Lambda (enables horizontal scaling)
  2. Database Connection Pooling: Reuse connections across requests
  3. Efficient Queries: Optimize SQL, use indexes, paginate results
  4. CDN Caching: Maximize CloudFront cache hit ratio
  5. Async Operations: Use queues for non-critical operations
  6. Monitoring: Proactive monitoring and alerting for limits