Skip to content
Singahi

Compliance · guide

ISO 27001 A.8.16: Monitoring Activities

104 min read

Share
On this page

Quick Reference: A.8.16 in 60 Seconds

QuestionAnswer
What is it?A control requiring organizations to monitor networks, systems, and applications for anomalous behavior and security events.
Why does it matter?You cannot defend what you cannot see. 60% of breaches are discovered by third parties, not the victim.
Minimum requirementEvent logging enabled on all critical systems + centralized log collection + review procedure + anomaly detection.
Audit red flagNo SIEM, logs only on local systems, no one reviewing logs, no alerting on critical events, no retention policy.
Quick winEnable centralized logging today. Forward Windows Event Logs to a collector. Forward Linux syslog to a central server.
Time to implement8–12 weeks for full monitoring capability. 2 weeks for basic log centralization.
Related controlsA.8.15 (Logging), A.8.17 (Clock Synchronization), A.5.7 (Threat Intelligence), A.5.24 (Incident Response), A.8.1 (Endpoints), A.8.5 (Authentication)

The 30-Second Architecture

┌─────────────────────────────────────────────────────────────┐
│  LAYER 4: RESPONSE (SOAR, Incident Response, Automation)    │
│  "Act on what you detect"                                   │
├─────────────────────────────────────────────────────────────┤
│  LAYER 3: ANALYSIS (SIEM, UEBA, Detection Engineering)      │
│  "Correlate, detect, and investigate"                       │
├─────────────────────────────────────────────────────────────┤
│  LAYER 2: AGGREGATION (Log Collection, Parsing, Normalization)│
│  "Collect everything in one place"                          │
├─────────────────────────────────────────────────────────────┤
│  LAYER 1: SOURCES (Endpoints, Network, Cloud, Apps, DBs)  │
│  "Generate logs from every asset"                           │
└─────────────────────────────────────────────────────────────┘

What the Standard Actually Requires

Figure · Process

What A.8.16 asks you to do

The 6 requirements of ISO 27001 A.8.16, monitoring activities, in order: monitoring strategy; monitoring tools; anomaly detection; event logging; log review; response to anomalies.
The 6 things the control expects. Each is expanded in the section below.

Figure · Matrix

Comparison: A.12.4.1, Event logging to A.12.4.5, Vulnerability

2022 VersionImplication
A.12.4.1, Event loggingMerged into A.8.16Logging is now part
A.12.4.2, ProtectionNow A.8.15Log protection
A.12.4.3, AdministratorMerged into A.8.16Privileged user monitoring
A.12.4.4, ClockNow A.8.17Time synchronization
A.12.4.5, VulnerabilityNow A.8.8Vulnerability management
Condensed from the table below, which carries the full detail for each cell.

ISO 27001:2022 A.8.16 Text

ISO 27001:2022 Annex A 8.16 asks organizations to monitor networks, systems, and applications for anomalous behavior and act on potential security incidents.

ISO 27002:2022 Implementation Guidance (Section 8.16)

ISO 27002 provides 6 implementation guidelines for 8.16:

  1. Monitoring strategy, Define what to monitor, how often, and with what tools. The strategy should cover networks, systems, and applications.
  2. Monitoring tools, Deploy tools to collect and analyze event data. Tools should support real-time alerting and historical investigation.
  3. Anomaly detection, Establish baselines and detect deviations from normal behavior. This includes both signature-based and behavioral detection.
  4. Event logging, Ensure that security-relevant events are logged. Logs should be protected from tampering and unauthorized access.
  5. Log review, Regularly review logs for signs of security incidents. Reviews should be documented and follow a defined procedure.
  6. Response to anomalies, Define how anomalies are evaluated and escalated. This links directly to A.5.24 (Incident Response) and A.5.25 (Assessment and Decision).

What Changed from ISO 27001:2013

2013 Version2022 VersionImplication
A.12.4.1, Event loggingMerged into A.8.16Logging is now part of monitoring, not separate
A.12.4.2, Protection of log informationNow A.8.15Log protection is a separate control
A.12.4.3, Administrator and operator logsMerged into A.8.16Privileged user monitoring is part of monitoring scope
A.12.4.4, Clock synchronizationNow A.8.17Time synchronization is separate but critical for monitoring
A.12.4.5, Vulnerability scanningNow A.8.8Vulnerability management is separate

Key implication: The 2022 version is more focused on active monitoring and response rather than passive log collection. The control now explicitly requires organizations to monitor networks, systems, and applications for anomalous behavior and act on potential security incidents. This means monitoring without response is non-compliant.

What Auditors Actually Check

Auditor ActionWhat They Want to See
Monitoring policyDocumented policy covering scope, tools, review frequency, and retention
SIEM/dashboardLive demonstration of centralized monitoring capability
Log source coverageEvidence that all critical systems forward logs to a central location
Alert configurationShow configured alerts for critical events (failed logins, privilege escalation, etc.)
Review recordsDocumented evidence that logs are reviewed regularly (daily/weekly)
Incident linkageEvidence that anomalies are escalated to incident response
Clock syncAll logs show consistent, synchronized timestamps (NTP configured)
Retention complianceLogs retained per policy and regulatory requirements
Tamper protectionLogs cannot be modified or deleted by administrators
Detection coverageMapping of detection rules to threats and risks in the risk register

The Monitoring Policy Minimum Requirements

Every auditor will ask for your monitoring policy. It must contain:

SectionMinimum Content
PurposeWhy monitoring exists, scope of information assets covered
ScopeNetworks, systems, applications, cloud services, endpoints, databases, physical access
RolesWho operates the SIEM, who reviews alerts, who responds to incidents
Log sourcesSpecific systems and event types that must be logged
CollectionHow logs are collected, forwarded, and centralized
RetentionRetention periods by log type (typically 90 days hot, 1 year warm, 3-7 years cold)
ReviewFrequency of review (real-time alerts, daily summary, weekly deep-dive, monthly trend)
Anomaly responseProcedure for evaluating and escalating anomalies
PrivacyPrivacy considerations for monitoring user activity, especially in EU/GDPR contexts
MetricsKPIs and performance indicators for the monitoring program

💡 Tip from Singahi: In our audits, we see that 45% of 8.16 non-conformities stem from a missing or inadequate monitoring policy. Organizations deploy Splunk or Sentinel but never document the process around it. The tool is not the control. The documented, repeatable process is the control.


The Monitoring Scope

Most guides stop at "networks, systems, and applications." That's incomplete. Here is the complete monitoring scope for a modern organization.

The 12 Monitoring Domains

DomainWhat to MonitorCritical EventsTypical Tools
1. NetworksTraffic patterns, flows, connections, bandwidth, protocolsUnusual outbound connections, C2 beaconing, lateral movementNetFlow collectors, IDS/IPS, NDR (Darktrace, Vectra)
2. Systems (Servers)OS events, process execution, file changes, loginsPrivileged access, unauthorized software, config changesOS native logging, OSQuery, Wazuh, Sysmon
3. ApplicationsLogin attempts, access patterns, API calls, errorsFailed authentication, injection attempts, abuse patternsApplication logs, APM tools (Datadog, New Relic), RASP
4. DatabasesQuery patterns, access logs, schema changes, connectionsUnusual queries, large exports, privilege escalationDAM (Imperva, Guardium), native audit logs, pgAudit
5. User ActivityLogins, file access, data downloads, email, webOff-hours access, bulk downloads, suspicious email rulesUEBA (Exabeam, Splunk UBA), DLP, CASB
6. EndpointsProcess execution, file changes, registry, networkMalware execution, credential dumping, USB insertionEDR (CrowdStrike, SentinelOne, Defender)
7. Cloud InfrastructureAPI calls, resource changes, access patterns, configurationsUnauthorized IAM changes, public bucket creation, unusual API volumeCloudTrail, GuardDuty, Azure Monitor, GCP SCC
8. APIsRequest volume, authentication, error rates, payloadsAbuse, injection, credential stuffing, excessive data exposureAPI gateways (Kong, Apigee), WAF logs, native API logs
9. Physical AccessBadge swipes, door alarms, CCTV, visitor logsTailgating, unauthorized area access, off-hours entryAccess control systems, physical security logs
10. Third-Party / Supply ChainVendor access logs, API integrations, data transfersUnauthorized vendor access, excessive data pull, integration abuseVendor portals, CASB, contractually mandated logs
11. Containers / KubernetesPod events, network policies, runtime behavior, admissionPrivilege escalation, suspicious pod creation, network policy violationsFalco, K8s audit logs, container runtime security
12. OT / ICSPLC events, SCADA logs, network segmentation, HMIUnauthorized engineering workstation access, protocol anomaliesOT-specific SIEMs, Nozomi Networks, Claroty

The Cloud-Native Expansion

Modern infrastructure has expanded the monitoring scope dramatically:

TraditionalCloud-Native AdditionWhy It Matters
Server logsContainer logs, pod events, K8s audit logsEphemeral infrastructure means logs must be captured in real-time
Network logsVPC Flow Logs, CloudWatch, Azure NSG logsCloud network traffic is invisible to traditional on-prem taps
Firewall logsCloud-native firewalls (AWS NFW, Azure FW), WAF logsApplication-layer attacks require Layer 7 visibility
Database logsCloud DB audit logs (RDS, Cosmos DB, BigQuery)Managed services have their own logging APIs
API logsAPI Gateway logs, Lambda execution logs, service mesh telemetryServerless functions generate logs you must explicitly configure
Physical accessCloud admin console access, IAM console login"Physical" access in cloud means console access

Edge Monitoring

Edge devices, IoT, retail POS, manufacturing sensors, remote offices, are often blind spots.

Edge ScenarioMonitoring ChallengeSolution
Remote office with limited bandwidthCannot forward all logs to central SIEMLightweight forwarders (Fluent Bit), edge aggregation, store-and-forward
IoT / OT networksProtocols are proprietary, devices are fragileProtocol gateways (Nozomi, Claroty), passive network monitoring
Mobile workforceDevices are off-network, inconsistent connectivityCloud-native EDR with offline caching, cloud SIEM ingestion
Air-gapped environmentsNo internet, no cloud SIEMOn-prem SIEM (Elastic, Splunk on-prem), data diode for export

💡 Tip from Singahi: During ISO 27001 Stage 2 audits, we see auditors specifically probe for cloud and container monitoring. If you run workloads in AWS or Azure but your monitoring policy only mentions "servers," expect a non-conformity. The scope must match your actual infrastructure.


The Monitoring Architecture

Design Principles

Defense in Depth for Monitoring

Don't rely on a single layer. Build monitoring depth:

┌─────────────────────────────────────────────────────────────┐
│  TIER 4: STRATEGIC (Threat Intelligence, Risk Dashboards) │
│  "What threats are coming? What's our exposure?"            │
├─────────────────────────────────────────────────────────────┤
│  TIER 3: OPERATIONAL (SOC, SIEM, Alerting, Response)      │
│  "What's happening right now? Should we respond?"           │
├─────────────────────────────────────────────────────────────┤
│  TIER 2: TACTICAL (Log Aggregation, Parsing, Storage)     │
│  "Collect, normalize, and store everything"                 │
├─────────────────────────────────────────────────────────────┤
│  TIER 1: TECHNICAL (Endpoint, Network, Cloud, App Logs)   │
│  "Generate telemetry from every asset"                      │
├─────────────────────────────────────────────────────────────┤
│  TIER 0: FOUNDATIONAL (Time Sync, Log Integrity, Policy)  │
│  "Without accurate time and integrity, monitoring is useless"│
└─────────────────────────────────────────────────────────────┘

Centralized vs. Distributed Monitoring

ModelDescriptionBest ForRisk
CentralizedAll logs flow to one SIEM. Single pane of glass.Small to mid-size orgs, single geographySingle point of failure, bandwidth saturation, data sovereignty issues
DistributedMultiple regional SIEMs or log aggregators.Multi-national orgs, data residency requirementsComplex correlation, tool sprawl, inconsistent detection
FederatedCentral analytics with regional collectors.Large enterprises, hybrid cloudHigher complexity, requires data normalization
HierarchicalEdge → Regional → Central tiers.OT/ICS environments, satellite officesLatency, potential data loss at tiers

Singahi recommendation: Start centralized. Most organizations under 1,000 employees should have one primary SIEM. Above that, consider federated with regional collectors. Above 10,000 employees or multi-national with data residency requirements, go hierarchical.

Real-Time vs. Batch Processing

Processing ModeLatencyUse CaseTool Examples
Real-time streaming< 1 secondCritical alerts, intrusion detection, automated responseKafka + Flink, Splunk RT, Sentinel streaming
Near-real-time1–5 minutesSecurity alerts, SIEM correlation, UEBAStandard SIEM ingestion, scheduled queries
Batch/hourly1–24 hoursCompliance reports, trend analysis, threat huntingScheduled reports, data warehouse queries
Deep archivalDays to weeksForensic investigation, long-term trend analysis, auditCold storage (S3 Glacier, Azure Archive)

Critical insight: Not everything needs to be real-time. Alerting should be real-time. Threat hunting can be near-real-time. Compliance reporting can be batch. Misjudging this creates alert fatigue and wastes infrastructure budget.

Agent-Based vs. Agentless Collection

ApproachHow It WorksProsConsBest For
Agent-basedSoftware installed on endpoint/serverRich data, local filtering, offline caching, response capabilityDeployment overhead, agent conflicts, resource consumptionEndpoints, critical servers, cloud VMs
AgentlessNetwork-based or API-based collectionNo deployment, no agent conflicts, lower overheadLimited data depth, network dependency, no offline capabilityNetwork devices, cloud services, legacy systems
HybridAgent-based for endpoints, agentless for network/cloudBest of both worldsComplexity, overhead, two toolchainsMost modern enterprises
Specific Collection MethodCategoryUse Case
Syslog / SNMP trapAgentlessNetwork devices, firewalls, printers
WMI / WinRMAgentlessWindows event collection without agent
Cloud API (CloudTrail, Azure Monitor)AgentlessCloud-native services
Database native auditAgentlessOracle, SQL Server, PostgreSQL audit
EDR agent (CrowdStrike, SentinelOne)Agent-basedEndpoint telemetry, threat detection, response
Filebeat / Fluent BitAgent-basedServer log forwarding, container log collection
OSQueryAgent-basedLive endpoint querying, fleet visibility
Packet capture (SPAN/TAP)AgentlessFull network traffic analysis, NDR

On-Premises vs. Cloud vs. Hybrid Monitoring

ArchitectureSIEM LocationCollection MethodComplexityoverhead Model
On-premisesSplunk, QRadar, Elastic on-premSyslog, agents, WEFHigh (self-managed infrastructure)CapEx + OpEx (hardware, licenses, staff)
Cloud-nativeSentinel, Chronicle, Sumo Logic, DatadogCloud APIs, cloud agentsLow (managed infrastructure)OpEx (per-user or per-ingestion licensing)
HybridCloud SIEM + on-prem collectorsHybrid agents, syslog forwardingMediumMixed (cloud license + on-prem hardware)
Multi-cloudSIEM in one cloud, collecting from othersCross-cloud IAM, API permissions, VPN/VPC peeringHighHighest (data egress, multi-cloud networking)

The Hybrid Reality: Most organizations are hybrid. They have on-prem Active Directory, AWS workloads, Azure SaaS, and GCP BigQuery. Your monitoring architecture must accommodate all three without creating blind spots.

Singahi's recommended hybrid architecture:

┌────────────────────────────────────────────────────────────────────┐
│                        CLOUD SIEM (Sentinel / Splunk Cloud)         │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────┐  ┌──────────┐ │
│  │ Analytics   │  │ Threat Intel│  │ SOAR        │  │ Reporting│ │
│  └─────────────┘  └─────────────┘  └─────────────┘  └──────────┘ │
└────────────────────────────────────────────────────────────────────┘
                              ▲
                              │ Encrypted TLS / VPN
┌────────────────────────────────────────────────────────────────────┐
│                    LOG AGGREGATION LAYER                            │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────┐  ┌──────────┐ │
│  │ Syslog NG   │  │ Kafka       │  │ Logstash    │  │ Fluentd  │ │
│  └─────────────┘  └─────────────┘  └─────────────┘  └──────────┘ │
└────────────────────────────────────────────────────────────────────┘
                              ▲
        ┌─────────────────────┼─────────────────────┐
        │                     │                     │
┌───────┴───────┐   ┌─────────┴─────────┐   ┌───────┴───────┐
│   ON-PREM     │   │      AWS           │   │    AZURE      │
│  AD, VMware   │   │  CloudTrail, VPC   │   │  Monitor, AD  │
│  Firewalls    │   │  GuardDuty, EKS    │   │  Sentinel,    │
│  Endpoints    │   │  CloudWatch, ALB   │   │  Defender     │
└───────────────┘   └───────────────────┘   └───────────────┘

💡 Tip from Singahi: The #1 architectural mistake we see is organizations buying Splunk Cloud and then discovering they cannot forward their on-prem firewall logs without a heavy forwarder or syslog gateway. Plan the data flow before you buy the license.


Log Sources & Collection

Operating System Logs

OSLog LocationCritical EventsCollection Method
WindowsSecurity (Event ID 4624/4625), System, Application, Sysmon, PowerShellFailed logins, privilege escalation, process creation, PowerShell executionWindows Event Forwarding (WEF), Winlogbeat, NXLog, Splunk UF
Linux/var/log/auth.log, /var/log/secure, /var/log/audit/audit.log, journaldSSH attempts, sudo usage, file integrity changes, SELinux denialsrsyslog, syslog-ng, auditd, Filebeat, Fluent Bit
macOSUnified Logging System (log), Open DirectoryLogin events, Gatekeeper blocks, file accessosquery, Splunk UF, syslog

Windows Event IDs Every Monitor Must Track:

Event IDDescriptionCriticalityWhy
4624Successful logonMediumBaseline activity; monitor for unusual patterns
4625Failed logonHighBrute force, credential stuffing, account reconnaissance
4648Explicit credential useHighPass-the-hash, lateral movement, runas abuse
4672Special privileges assignedHighAdmin login, privilege escalation
4688Process creationMediumMalware execution, living-off-the-land binaries
4697Service installedHighPersistence via malicious service
4698Scheduled task createdHighPersistence mechanism
4720User account createdHighAccount creation (insider threat, attacker)
4726User account deletedHighAccount deletion (cover-up)
4728Member added to security groupHighPrivilege escalation, group membership abuse
4732Member added to local groupHighLocal admin addition
4768Kerberos TGT requestedMediumKerberos abuse, AS-REP Roasting
4769Kerberos service ticket requestedMediumGolden Ticket, Silver Ticket detection
4771Kerberos pre-auth failedHighKerberoasting, password guessing
4776NTLM authenticationMediumPass-the-hash, downgrade attacks
5136Directory service object modifiedHighAD object modification, DCSync preparation
7045Service installed (System)HighMalware persistence

Linux Audit Rules Every Organization Should Deploy:

## Monitor user/group modifications
-w /etc/passwd -p wa -k identity_changes
-w /etc/group -p wa -k identity_changes
-w /etc/shadow -p wa -k identity_changes

## Monitor sudoers
-w /etc/sudoers -p wa -k sudoers_changes
-w /etc/sudoers.d -p wa -k sudoers_changes

## Monitor SSH configuration
-w /etc/ssh/sshd_config -p wa -k ssh_config_changes

## Monitor privilege escalation
-a always,exit -F arch=b64 -S setuid -S setgid -k privilege_escalation

## Monitor file integrity critical paths
-w /etc/cron.d -p wa -k cron_changes
-w /etc/cron.daily -p wa -k cron_changes
-w /var/spool/cron -p wa -k cron_changes

## Monitor suspicious command execution
-a always,exit -F arch=b64 -S execve -C uid!=euid -k privilege_escalation

Application Logs

Application TypeWhat to LogFormatCollection
Web serversHTTP requests, errors, authentication, access deniedApache Combined Log, Nginx, IIS W3CFilebeat, Fluent Bit, direct syslog
DatabasesConnections, queries, errors, privilege changesNative audit log, query log, general logDatabase audit agents, JDBC/ODBC forwarding
Custom applicationsAuthentication, authorization, business events, errorsStructured JSON (preferred) or key-valueApplication logging libraries (log4j, winston, zap) → forwarder
ERP/CRMUser logins, record access, configuration changes, exportsProprietary or database-backedAPI extraction, database trigger logging
EmailSend/receive, authentication, attachment types, forwarding rulesSMTP logs, mailbox audit logsExchange Admin Center, Office 365 audit, Google Workspace audit

Application Logging Best Practices (The 5-Field Rule):

Every application log entry should contain at minimum:

  1. Timestamp (ISO 8601, UTC, millisecond precision)
  2. Severity (DEBUG, INFO, WARN, ERROR, CRITICAL)
  3. Event type (auth.success, auth.failure, data.access, config.change)
  4. Actor (user ID, service account, IP address, session ID)
  5. Outcome (success, failure, blocked, allowed, with detail)

Example:

{
  "timestamp": "2026-06-15T05:19:50.173Z",
  "severity": "WARN",
  "event_type": "auth.failure",
  "actor": {
    "user_id": "user@company.com",
    "ip_address": "203.0.113.45",
    "session_id": "sess_abc123"
  },
  "outcome": "failure",
  "detail": {
    "reason": "invalid_password",
    "mfa_required": true,
    "attempt_count": 3
  }
}

Network Logs

Log TypeWhat It CapturesKey ValueCollection
Firewall logsAllowed/blocked connections, NAT, VPNPerimeter control visibilitySyslog, API (cloud firewalls), CEF
DNS logsQueries, responses, NXDOMAIN, DGA patternsMalware C2 detection, data exfiltration via DNS tunnelingDNS server logs (BIND, Windows DNS), DNS firewalls (Infoblox), Pi-hole
Proxy logsWeb requests, URL filtering, user identityWeb browsing behavior, shadow IT, malicious downloadsSquid, BlueCoat, Zscaler, Netskope
VPN logsConnections, authentication, IP assignment, durationRemote access abuse, compromised credential useVPN concentrator logs, RADIUS logs
Load balancer logsHTTP requests, TLS handshake, backend health, source IPDDoS detection, application abuse, geo-anomaliesALB/NLB logs (AWS), Application Gateway (Azure), NGINX, HAProxy
WAF logsBlocked requests, rule triggers, IP reputationApplication attack detection (SQLi, XSS, LFI)AWS WAF, Azure WAF, Cloudflare, Imperva, F5 ASM

Authentication & Identity Logs

SourceCritical EventsDetection Value
Active DirectoryLogon/logoff, group changes, password changes, Kerberos eventsAccount compromise, privilege escalation, lateral movement
Azure AD / Entra IDSign-in logs, audit logs, risk events, MFA failuresCloud identity compromise, conditional policy bypass
OktaUser sessions, MFA events, app assignments, suspicious activitySSO compromise, session hijacking
AWS IAMConsole login, API calls, IAM policy changes, role assumptionCloud account compromise, privilege escalation
SAML / OIDC providersAssertion issuance, token validation, logoutFederation abuse, token replay
RADIUS / TACACS+Network device authentication, authorization, accountingNetwork infrastructure compromise

Database Audit Logs

DatabaseAudit MechanismCritical EventsTools
Microsoft SQL ServerSQL Server Audit, Extended EventsLogin failures, privilege changes, schema modifications, sensitive table accessSQL Server Audit, Splunk DB Connect
PostgreSQLpgaudit extension, log_statementDDL changes, DML on sensitive tables, connection attemptspgaudit, Filebeat
MySQLGeneral query log, audit log pluginLogin failures, privilege changes, slow queriesMySQL Enterprise Audit, Percona Audit
OracleUnified Audit Policy, Fine-Grained AuditingPrivileged user actions, sensitive data access, DBA activityOracle Audit Vault, Imperva
MongoDBAudit log (Enterprise), profilerAuthentication, authorization, CRUD on sensitive collectionsMongoDB Atlas logs, Filebeat
ElasticsearchAudit log (Security API)Index access, cluster changes, privilege changesElastic Security
SnowflakeAccount usage views, access historyQuery history, login history, warehouse usageSnowflake's native views, external ingestion
AWS RDSCloudTrail (API), parameter group audit logs, Enhanced MonitoringInstance modifications, snapshot access, performance anomaliesCloudTrail, RDS console

Cloud Audit Logs

Cloud ServiceLog SourceCritical EventsNotes
AWS CloudTrailManagement events, data events, insight eventsIAM changes, root account use, bucket policy changes, unauthorized API callsEnable in all regions. Enable data events for S3 and Lambda.
AWS GuardDutyThreat detection findingsCredential compromise, EC2 malware, data exfiltration, reconnaissanceNative ML-based. Feed findings into SIEM.
AWS VPC Flow LogsIP traffic flowUnusual connections, data exfiltration, lateral movementAnalyzing at scale requires dedicated tools or Athena.
AWS ConfigResource configuration changesNon-compliant resource changes, drift detectionCompliance-focused, not real-time threat detection.
Azure Activity LogSubscription-level eventsResource creation, RBAC changes, policy changesFeed into Log Analytics workspace.
Azure Sign-in LogsAuthentication eventsRisky sign-ins, MFA failures, conditional access blocksRequires Azure AD P1/P2.
Azure Defender for CloudSecurity alerts, recommendationsVM malware, container security, SQL threat detectionNative security monitoring.
GCP Cloud Audit LogsAdmin Activity, Data Access, System Event, Policy DeniedIAM changes, resource access, policy violationsAutomatically enabled. Feed into BigQuery or SIEM.
GCP Security Command CenterFindings, assets, vulnerabilitiesMisconfigurations, threats, sensitive data exposurePremium tier adds threat detection.
OCI Audit LogsTenancy-level eventsIAM changes, resource access, policy changesFeed into Logging Analytics or external SIEM.

Kubernetes & Container Logs

Log SourceWhat It CapturesCritical EventsCollection
K8s Audit LogsAPI server requestsPod creation, RBAC changes, secret access, exec into podsEnable --audit-log-path on API server. Fluent Bit/Fluentd collection.
K8s Event LogsResource events (pod scheduling, failures)CrashLoopBackOff, ImagePullBackOff, OOMKilledkubectl get events or collected via event exporter
Container Runtime (containerd, CRI-O)Container lifecycle, image pulls, execPrivileged container creation, host namespace usage, suspicious image pullsContainer runtime logs (often via journald)
FalcoRuntime security eventsSyscall anomalies, privileged container escapes, sensitive file accessFalco sidecar or daemonset, output to stdout or gRPC
Service Mesh (Istio, Linkerd)mTLS traffic, request metrics, access policiesUnauthorized service communication, policy violations, anomalous trafficIstio Envoy access logs, Prometheus metrics
Admission ControllersPolicy enforcement (OPA, Kyverno)Rejected deployments, policy violations, mutated resourcesController logs, often in K8s audit log

Physical Access Logs

SystemData CapturedSecurity ValueIntegration
Badge access systems (HID, Lenel, Genetec)Badge ID, door, timestamp, access granted/deniedUnauthorized area access, tailgating, off-hours entrySIEM via syslog or API; correlation with logical access
Visitor management (Envoy, Proxyclick)Visitor name, host, check-in/check-out, photoUnescorted visitors, overstays, badge handoffAPI integration, manual review
CCTV / VMS (Genetec, Milestone, Avigilon)Video streams, motion detection, alarm eventsVisual verification of access events, investigation supportTypically not integrated into SIEM; alarm events can be
Elevator accessFloor access by badgeRestricted floor access, after-hours movementOften integrated with badge system

API Gateway Logs

GatewayLog FormatCritical DataSecurity Use
KongAccess logs, error logs, plugin logsConsumer ID, route, latency, status code, request/response sizeAPI abuse detection, rate limit violations, authentication failures
AWS API GatewayCloudWatch Logs, Kinesis Data FirehoseRequest ID, IP, user agent, integration latency, WAF integrationInjection detection, unusual request patterns, DDoS
Azure API ManagementApplication Insights, Event HubRequest/response bodies (optional), backend response, subscription keyUnauthorized API calls, data exfiltration via API
ApigeeAnalytics, Edge message logsDeveloper app, product, resource, fault codesAPI monetization abuse, developer policy violations
NGINX / EnvoyStandard access logs + custom variablesUpstream response, TLS version, SNI, rate limit statusMicroservices security, service mesh monitoring

💡 Tip from Singahi: The most common log source gap is DNS. Organizations monitor firewalls and endpoints but ignore DNS. 80% of malware uses DNS for C2 communication. If you're not logging DNS queries, you're blind to the majority of modern malware traffic. Enable DNS logging today, it's often free with your existing DNS infrastructure.


Log Management & Centralization

Syslog, Syslog-NG, and Rsyslog

The syslog protocol (RFC 5424) is the backbone of log collection. Understand the differences:

DaemonProsConsBest For
rsyslogDefault on most Linux distros, high performance, modularComplex configuration syntax, debugging can be painfulStandard Linux log aggregation, medium scale
syslog-ngCleaner config, better parsing, log routing, patternDBNot default on most distros, learning curveLarge-scale Linux environments, complex routing needs
nxlogCross-platform (Windows + Linux), structured config, powerfulEnterprise edition for advanced featuresWindows-heavy environments, unified agent
Vector (by Datadog)High performance, VRL (Vector Remap Language), observabilityNewer, smaller communityCloud-native, high-volume environments

Sample rsyslog Configuration for Centralized Collection:

## /etc/rsyslog.conf
## Enable TCP reception
module(load="imtcp")
input(type="imtcp" port="514")

## Enable UDP reception (fallback only)
module(load="imudp")
input(type="imudp" port="514")

## Template for structured storage
*.* ?RemoteLogs
& stop

## Forward everything to central SIEM over TLS
*.* @@central-siem.company.com:6514;RSYSLOG_SyslogProtocol23Format

Sample syslog-ng Configuration with Parsing:

## /etc/syslog-ng/syslog-ng.conf
source s_network {
    network(
        transport("tcp")
        port(514)
    );
};

destination d_central {
    syslog("central-siem.company.com"
        transport("tls")
        port(6514)
        tls(
            ca_dir("/etc/syslog-ng/ca.d")
            key_file("/etc/syslog-ng/client.key")
            cert_file("/etc/syslog-ng/client.crt")
        )
    );
};

## Parse and route Apache logs
filter f_apache {
    program("apache");
};

destination d_apache_parsed {
    file("/var/log/apache/structured.log"
    );
};

log {
    source(s_network);
    filter(f_apache);
    destination(d_apache_parsed);
    destination(d_central);
};

Windows Event Forwarding (WEF)

WEF is the native, agentless way to collect Windows events. It uses WinRM and is built into Windows Server.

Architecture:

Domain Controllers / Workstations (Source)
           ↓ (WinRM, HTTPS or HTTP with Kerberos)
    Windows Event Collector (WEC) Server
           ↓ (rsyslog, Splunk UF, or direct to SIEM)
           SIEM

Why WEF matters:

  • No agent installation on endpoints (agentless)
  • Uses Active Directory for authentication
  • Supports both push (source-initiated) and pull (collector-initiated) subscriptions
  • Can filter events at the source, reducing bandwidth

WEF Subscription Types:

TypeDescriptionUse Case
Source-initiatedEndpoints forward to collector based on AD groupLarge environments, roaming laptops
Collector-initiatedCollector pulls from specific serversSmall environments, specific server monitoring

WEF Limitations:

  • Only works for Windows events (not Sysmon, not custom application logs)
  • Requires Active Directory
  • No native encryption without HTTPS configuration
  • Limited parsing capability (events are XML)

Best practice: Use WEF for baseline Windows Security events. Use an agent (Winlogbeat, Splunk UF, CrowdStrike) for rich endpoint telemetry including Sysmon, PowerShell, and custom application logs.

Log Shippers: Fluentd, Fluent Bit, Logstash, Vector

Modern log architecture uses lightweight log shippers at the edge.

ShipperLanguageResource UsageBest ForNotable Features
FluentdRuby (JVM)Medium (40MB+ RAM)Complex parsing, rich plugin ecosystem500+ plugins, tag-based routing
Fluent BitCVery low (<1MB RAM)Containers, edge devices, high-volumeK8s-native, can replace Fluentd in most cases
LogstashJava (JVM)High (1GB+ RAM)Heavy transformation, Elastic integrationPowerful filter plugins, Beats input, high overhead
VectorRustLowHigh-performance, cloud-nativeVRL language, observability, multiple sinks
FilebeatGoLowFile-based log collection, Elastic stackLightweight, backpressure handling, modules
PromtailGoLowGrafana Loki stackK8s service discovery, label extraction

Fluent Bit Configuration Example (Kubernetes + CloudWatch + SIEM):

## fluent-bit-config.yaml
apiVersion: v1
kind: ConfigMap
metadata:
  name: fluent-bit-config
data:
  fluent-bit.conf: |
    [SERVICE]
        Flush         1
        Log_Level     info
        Daemon        off
        Parsers_File  parsers.conf
    [INPUT]
        Name              tail
        Tag               kube.*
        Path              /var/log/containers/*.log
        Parser            docker
        DB                /var/log/flb_kube.db
        Mem_Buf_Limit     50MB
    [FILTER]
        Name                kubernetes
        Match               kube.*
        Kube_URL            https://kubernetes.default.svc:443
        Kube_CA_File        /var/run/secrets/kubernetes.io/serviceaccount/ca.crt
        Kube_Token_File     /var/run/secrets/kubernetes.io/serviceaccount/token
    [OUTPUT]
        Name            cloudwatch_logs
        Match           kube.*
        region          us-east-1
        log_group_name  /eks/application-logs
    [OUTPUT]
        Name            syslog
        Match           kube.*
        Host            siem.company.com
        Port            514
        Mode            tcp
        Syslog_Format   rfc5424

Log Parsing and Normalization

Raw logs are useless without parsing. The goal is to convert unstructured or semi-structured logs into a common schema.

Common Log Schemas:

SchemaOrganizationUse CaseEcosystem
ECS (Elastic Common Schema)ElasticElastic Stack, open standardElastic, many third-party tools
CIM (Common Information Model)SplunkSplunk Enterprise SecuritySplunk-specific
ASIM (Advanced Security Information Model)MicrosoftMicrosoft SentinelAzure/Sentinel-specific
CEF (Common Event Format)ArcSightLegacy but widely supportedArcSight, many firewalls, legacy tools
LEEF (Log Event Extended Format)IBMQRadar integrationIBM QRadar
JSON (custom)YouFlexibility, human readabilityModern applications, custom parsers
OTel (OpenTelemetry)CNCFObservability, metrics + logs + tracesCloud-native, vendor-neutral

Parsing Example: Apache Combined Log to ECS:

Raw: 192.168.1.1 - - [15/Jun/2026:05:19:50 +0000] "GET /api/users HTTP/1.1" 200 1234 "https://app.company.com" "Mozilla/5.0"

ECS Parsed:
- source.ip: 192.168.1.1
- http.request.method: GET
- url.path: /api/users
- http.response.status_code: 200
- http.response.body.bytes: 1234
- http.request.referrer: https://app.company.com
- user_agent.original: Mozilla/5.0
- event.category: web
- event.dataset: apache.access

Normalization Pipeline:

┌─────────────┐    ┌─────────────┐    ┌─────────────┐    ┌─────────────┐
│  Raw Log    │ →  │   Parse     │ →  │  Enrich     │ →  │  Normalize  │
│  (Apache,   │    │  (Regex,    │    │  (GeoIP,    │    │  (ECS, CIM,│
│  Windows,   │    │  Grok,      │    │  Threat     │    │  ASIM)      │
│  JSON)      │    │  Dissect)   │    │  Intel, AD) │    │             │
└─────────────┘    └─────────────┘    └─────────────┘    └─────────────┘

Log Aggregation Architecture Patterns

Pattern 1: Direct to SIEM (Small Orgs)

Endpoints → Splunk UF / Filebeat → Splunk / Elastic
  • Simple, low latency
  • No intermediate buffering
  • Risk: SIEM outage = data loss

Pattern 2: Message Queue Buffering (Mid-Size)

Endpoints → Fluent Bit → Kafka → Logstash → Elastic
  • Kafka provides buffering and replay
  • Decouples collection from analysis
  • Better for high-volume or bursty environments

Pattern 3: Cloud-Native Pipeline (Cloud-First)

CloudWatch / Azure Monitor / GCP Logging → Data Firehose / Event Hub → S3 / Data Lake → Athena / Spark → SIEM
  • efficient for high volume
  • Enables long-term storage and batch analytics
  • Higher latency for real-time alerting

Pattern 4: Multi-Cloud Federation (Enterprise)

AWS CloudTrail → Kinesis → Lambda (normalize) → Kafka
Azure Activity Logs → Event Hub → Function (normalize) → Kafka
On-Prem syslog → syslog-ng → Kafka
Kafka → Splunk / Sentinel / Elastic (central SIEM)
  • Unified schema across all sources
  • Complex to build and maintain
  • Requires dedicated data engineering team

💡 Tip from Singahi: We recommend Pattern 2 for organizations with 100–5,000 employees and Pattern 4 for organizations above 5,000 employees or with multi-cloud requirements. Pattern 1 is acceptable for under 100 employees but will create pain as you scale.


SIEM Deep Dive

A SIEM is the brain of your monitoring operation. Choosing the wrong one is a multi-year mistake. This section covers the major platforms in forensic detail.

Splunk Enterprise / Splunk Cloud

AttributeDetail
ArchitectureProprietary SPL (Search Processing Language), indexed data, schema-on-read
DeploymentOn-prem (Enterprise), SaaS (Cloud), or hybrid (Cloud with on-prem forwarders)
IngestionHeavy Forwarders, Universal Forwarders, HTTP Event Collector (HEC), syslog, APIs
StrengthsMost powerful search language, largest app ecosystem, mature SOAR (Splunk SOAR), strong community
Weaknessespremium-tier at scale, SPL learning curve, resource-heavy infrastructure
Best forLarge enterprises, complex correlation, organizations with existing Splunk investment

Splunk Enterprise Security (ES):

  • Pre-built correlation searches for common use cases
  • Risk scoring framework
  • Notable event management
  • Threat intelligence framework (STIX/TAXII)
  • Incident Review dashboard
  • overhead: Additional license on top of Splunk base

Splunk SOAR:

  • Playbook automation
  • 300+ integrations
  • Case management
  • overhead: Additional license, but essential for SOC efficiency

Splunk for ISO 27001:

  • Pre-built compliance dashboard for ISO 27001
  • Data models for Authentication, Change, Endpoint, Network
  • Ability to map events to Annex A controls

IBM QRadar

AttributeDetail
ArchitectureProprietary Ariel database, normalized event format (QRadar DSM)
DeploymentOn-prem (QRadar SIEM), SaaS (QRadar on Cloud), or QRadar SIEM Virtual
IngestionSyslog, APIs, QRadar WinCollect, QRadar Log Source Management
licensingPer Events Per Second (EPS) or Flows Per Minute (FPM). Growing-company friendly.
StrengthsExcellent out-of-box rules, strong network behavior analytics, good for regulated industries (finance, government)
WeaknessesAging UI, slower search for large datasets, smaller app ecosystem than Splunk
Best forFinancial services, government, organizations with existing IBM security stack

QRadar Differentiators:

  • QRadar Network Insights: Deep packet inspection without decryption
  • QRadar User Behavior Analytics (UBA): Free add-on for user anomaly detection
  • QRadar Vulnerability Manager: Integration with vulnerability data for risk-based prioritization
  • QRadar Advisor with Watson: AI-assisted threat investigation (limited value in practice)

Microsoft Sentinel

AttributeDetail
ArchitectureAzure-native, built on Log Analytics workspace, KQL (Kusto Query Language)
DeploymentCloud-only (Azure). Collectors for on-prem, multi-cloud, and SaaS.
IngestionAzure Monitor Agent, Log Analytics API, Azure Event Hub, third-party connectors (50+), CEF/syslog via Log Analytics agent
StrengthsNative Azure integration, excellent Microsoft 365/Defender integration, KQL is powerful, SOAR built-in (Logic Apps), efficient for Microsoft-heavy orgs
WeaknessesLess mature for non-Microsoft environments, KQL learning curve, can become premium-tier at scale without commitment tiers
Best forMicrosoft-heavy organizations (Azure, M365, Defender), growing companies, cloud-first companies

Sentinel Key Features:

  • Microsoft 365 Defender Integration: Automatic ingestion of Defender for Endpoint, Identity, Cloud Apps, and Office 365 alerts
  • Built-in SOAR: Azure Logic Apps playbooks for automation
  • Microsoft Threat Intelligence: Free integration with Microsoft TI feeds
  • Hunting Queries: Community-driven KQL queries (GitHub: Azure/Azure-Sentinel)
  • Fusion ML: Multi-stage attack detection across signals
  • UEBA: Built-in (requires Azure AD P2)

Sentinel licensing Reality:

  • A 500-person company with Microsoft 365 E5 gets significant data ingestion included

Google Chronicle

AttributeDetail
ArchitectureGoogle-native, hyper-scale backend, YARA-L detection language
DeploymentCloud-only (Google Cloud). Collectors for multi-cloud and on-prem.
IngestionChronicle Forwarder, Google Cloud Logging, APIs, CEF/syslog
licensingPer GB/day or per user/year. Marketed as "unlimited retention" with per-user licensing.
StrengthsUnlimited retention (differentiating feature), extremely fast search over years of data, YARA-L for detection, native VirusTotal integration, Mandiant threat intel
WeaknessesSmaller ecosystem, newer product, less third-party integration than Splunk/Sentinel, requires Google Cloud adoption
Best forSecurity teams prioritizing threat hunting over real-time SOC, organizations with Google Cloud presence, Mandiant customers

Chronicle Unique Value:

  • Unlimited Retention: All data retained at no extra overhead. This is a paradigm shift for forensic investigations.
  • Asset-Entity Graph: Automatic entity extraction (users, devices, IPs) and relationship mapping
  • Curated Detections: Google/Mandiant-authored YARA-L rules
  • VirusTotal Enterprise: Included for file/hash investigation

Elastic Security (Elastic Stack)

AttributeDetail
ArchitectureOpen-source Elasticsearch + Kibana + Elastic Agent/Beats. Elastic Security is the SIEM UI.
DeploymentSelf-managed (on-prem/cloud VMs), Elastic Cloud (managed), or Elastic Cloud on Kubernetes
IngestionElastic Agent (unified), Filebeat, Winlogbeat, Packetbeat, Auditbeat, Metricbeat, custom integrations
licensingSelf-managed: free (open source) or Elastic license for advanced features. Elastic Cloud: resource-based.
StrengthsBest open-source SIEM, extremely fast search, flexible schema (ECS), growing detection rule library, no per-GB licensing for self-managed
WeaknessesSelf-managed requires expertise, less polished SOAR than competitors, smaller enterprise support ecosystem
Best foroverhead-conscious organizations, tech-savvy teams, open-source advocates, high-volume data environments

Elastic Security Components:

  • Elastic Agent: Unified agent replacing Beats
  • Fleet: Centralized agent management
  • Detection Engine: Rule-based detection with KQL + EQL (Event Query Language)
  • Cases: Built-in case management
  • Osquery Manager: Live endpoint querying
  • Threat Intelligence: STIX/TAXII integration, custom IOC lists
  • ML Jobs: Anomaly detection for rare events, unusual traffic, beaconing

Elastic licensing (Self-Managed):

  • Open source (free): Core SIEM functionality
  • Key advantage: You pay for infrastructure, not per GB. High-volume environments can be dramatically cheaper.

LogRhythm

AttributeDetail
ArchitectureProprietary AI Engine, centralized log management, integrated UEBA
DeploymentOn-prem (XM/PM) or SaaS (LogRhythm Cloud)
IngestionSyslog, agents, APIs, Beats, CEF
licensingPer node or per GB. Growing-company focused.
StrengthsStrong integrated UEBA, good case management out-of-box, analyst-friendly UI, fast time-to-value
WeaknessesSmaller ecosystem, less cloud-native, slower innovation pace
Best forGrowing companies wanting integrated UEBA without buying separate tools

Sumo Logic

AttributeDetail
ArchitectureCloud-native, multi-tenant, proprietary query language
DeploymentCloud-only (SaaS). Collectors for on-prem and cloud.
IngestionInstalled Collectors, Hosted Collectors, APIs, CloudWatch, Azure Monitor, GCP
licensingPer GB/day ingested. Freemium tier available.
StrengthsPure cloud-native, good for DevOps/SRE teams, strong log analytics, PCI compliance features
WeaknessesLess security-focused than Splunk/Sentinel, smaller SOAR capability, not ideal for pure SOC use
Best forDevOps-heavy organizations, cloud-native companies, hybrid security + observability use case

Datadog Security Monitoring

AttributeDetail
ArchitectureCloud-native, unified observability + security platform
DeploymentCloud-only (SaaS). Agents for infrastructure.
IngestionDatadog Agent, Cloud integrations, APIs, Log Forwarding
licensingPer host/month (infrastructure) + per GB/day (logs) + per host/month (security).
StrengthsStrong observability integration, Cloud Security Posture Management (CSPM) included, Cloud Workload Security (CWS) for runtime threat detection, SIEM is improving rapidly
WeaknessesSecurity SIEM is newer and less mature than dedicated SIEMs, licensing can escalate quickly with full stack
Best forOrganizations already using Datadog for APM/infrastructure monitoring, cloud-native DevOps teams

Datadog Security Products:

  • Security Monitoring: SIEM capability (log analysis, detection rules)
  • Cloud SIEM: Full SIEM with dashboards, case management
  • Cloud Security Posture Management (CSPM): Misconfiguration detection
  • Cloud Workload Security (CWS): Runtime threat detection (Falco-like)
  • Application Security Management (ASM): In-app threat detection (RASP-like)

New Relic Security

AttributeDetail
ArchitectureObservability-first, security add-on
DeploymentCloud-only (SaaS)
licensingPer GB/day + per user
StrengthsStrong observability foundation, good for application security
WeaknessesSecurity SIEM is immature, not a primary SOC platform
Best forOrganizations already invested in New Relic wanting to add security context

Wazuh (Open Source)

AttributeDetail
ArchitectureOpen-source SIEM + XDR + CSPM. Fork of OSSEC.
DeploymentSelf-managed (Linux server + agents). Free.
IngestionWazuh agents, syslog, cloud APIs, Docker integration
licensingFree (open source). Commercial support available.
StrengthsFree, includes EDR-like capability (file integrity, rootkit detection, active response), integrates with Elastic/Kibana, good for compliance
WeaknessesRequires significant self-management, smaller detection rule library, limited SOAR, scalability challenges beyond 10,000 agents
Best forBudget-constrained organizations, small SOC teams, compliance-first monitoring

Wazuh Capabilities:

  • Log Analysis: Collects and analyzes logs from endpoints, network devices, applications
  • File Integrity Monitoring (FIM): Detects file changes in real-time
  • Malware Detection: Rootkit detection, anomaly detection
  • Active Response: Automated blocking of IPs, disabling accounts
  • Vulnerability Detection: Scans endpoints for CVEs
  • CSPM: Cloud security posture for AWS, Azure, GCP
  • Container Security: Docker and Kubernetes monitoring

Graylog (Open Source)

AttributeDetail
ArchitectureOpen-source log management. Built on MongoDB + Elasticsearch/OpenSearch.
DeploymentSelf-managed. Enterprise version adds features.
IngestionSyslog, GELF, Beats, APIs, AWS Kinesis, Kafka
licensingOpen source (free). Enterprise license for alerts, archiving, audit.
StrengthsSimple, fast search, good syslog collection, easy to deploy, lightweight
WeaknessesNot a full SIEM (limited correlation, no UEBA, no SOAR), requires self-management
Best forLog aggregation and search, smaller teams, syslog-heavy environments

SIEM Comparison by Organization Size

Organization SizeRecommended SIEMBudget RangeWhy
Cloud-native / DevOpsDatadog Security, Chronicle, or ElasticVariesPrioritize observability integration and cloud-native architecture.

SIEM licensing Models Explained

ModelDescriptionWhen It's CheaperWhen It's premium-tier
Per GB ingestedSplunk, Sentinel, Datadog logsLow volume, selective ingestionHigh volume, verbose applications
Per EPS (Events Per Second)QRadarSteady, predictable event ratesBursty traffic, event storms
Per user/hostDatadog, Chronicle, some Sentinel tiersFixed infrastructure sizeRapid growth, scaling unpredictably
Per node/deviceLogRhythmSmall number of high-value assetsLarge, distributed endpoint population
Infrastructure-basedSelf-managed Elastic, WazuhHigh volume, dedicated teamSmall volume, limited ops expertise
Freemium / tieredSumo Logic, SentinelStartups, gradual growthCrossing tier thresholds unexpectedly

overhead Optimization Strategies:

  1. Filter before ingest: Drop debug logs, health checks, and duplicate events at the collector
  2. Use hot/warm/cold tiers: Only keep recent data in premium-tier hot storage
  3. Selective parsing: Not every field needs to be indexed
  4. Sample high-volume logs: Application performance metrics can be sampled 1:100
  5. Dedicated ingestion windows: Batch non-critical logs during off-peak hours
  6. Negotiate commit tiers: Sentinel and Splunk offer significant discounts for volume commitments

Security Monitoring Use Cases

This is where most guides stay generic. Here are specific, actionable detection use cases with exact detection logic.

Use Case 1: Account Compromise

Scenario: Attacker obtains credentials via phishing or credential stuffing and logs into a corporate account.

Detection Logic:

SignalDetection RuleSeverity
Impossible travelLogin from two geographically distant locations within impossible time windowHigh
Off-hours loginFirst login outside business hours (e.g., 10 PM–5 AM) for this user in 90 daysMedium
New device / browserLogin from a device or browser never seen for this userMedium
MFA bypassSuccessful login after multiple MFA failuresHigh
Password sprayMultiple failed logins across many accounts from single IPHigh
Suspicious email ruleInbox rule created to forward all email to external addressCritical
Login after breach dataPassword found in known breach database (via HIBP or similar)Medium

Example Splunk SPL:

| tstats `summariesonly` count from datamodel=Authentication by _time, src_user, src_ip, action
| where action="failure"
| stats count by src_user, src_ip
| where count > 5
| lookup geo_ip src_ip as src_ip
| where distance(src_lat, src_lon, user_home_lat, user_home_lon) > 500 AND time_window < 2h
| eval severity="high"

Use Case 2: Privilege Escalation

Scenario: Attacker or insider elevates privileges to gain unauthorized access.

Detection Logic:

SignalDetection RuleSeverity
Local admin additionUser added to local Administrators group (Event ID 4732)High
Domain admin additionUser added to Domain Admins (Event ID 4728)Critical
sudo abuseUnusual sudo commands (e.g., sudo su, sudo bash) by non-admin userHigh
IAM policy changeIAM policy attached to user or role allowing new permissions (CloudTrail)High
Azure AD role changePrivileged role assigned (e.g., Global Administrator)Critical
Kubernetes RBAC changeClusterRoleBinding or RoleBinding created with elevated privilegesHigh
Pass-the-hashEvent ID 4648 with NTLM and suspicious process (e.g., mimikatz indicators)Critical
KerberoastingService ticket request (Event ID 4769) with weak encryption (RC4)High
DCSyncDirectory replication request from non-DC account (Event ID 4662)Critical

Example Sigma Rule (Privilege Escalation, Local Admin):

title: User Added to Local Administrators
logsource:
  product: windows
  service: security
detection:
  selection:
    EventID: 4732
    TargetUserName: Administrators
  condition: selection
level: high

Use Case 3: Lateral Movement

Scenario: Attacker moves from compromised endpoint to other systems within the network.

Detection Logic:

SignalDetection RuleSeverity
RDP hopUser RDPs to Server A, then RDPs from Server A to Server BHigh
PSExec executionSysmon Event ID 1 with Image containing "psexec" or "psexesvc"High
WMI executionSysmon Event ID 1 with ParentImage containing "wmic"High
PowerShell remotingWinRM or PowerShell remoting connections from workstationsMedium
SSH lateral movementSSH from server to server (unusual for this source)Medium
Service creation for lateral movementService created on remote host (Event ID 4697) from unusual sourceHigh
Pass-the-ticketKerberos TGS request with abnormal ticket lifetime or encryptionHigh

Use Case 4: Data Exfiltration

Scenario: Sensitive data is transferred out of the organization.

Detection Logic:

SignalDetection RuleSeverity
Large outbound transferOutbound data volume > 3x baseline for this user/systemHigh
Unusual cloud uploadBulk upload to personal cloud storage (OneDrive personal, Google Drive, Dropbox)High
Database bulk exportSELECT query returning > 10,000 rows on sensitive tableHigh
Email with large attachmentEmail sent to external domain with attachment > 50MBMedium
USB mass storageUSB device with high read/write volume on sensitive file serverHigh
DNS tunnelingHigh volume of DNS queries to unusual domains, or TXT records with large payloadsHigh
HTTPS to suspicious domainLarge HTTPS transfer to newly registered or DGA domainHigh
Print to PDF abuseMass printing to PDF printer on sensitive document repositoryMedium

Example KQL (Sentinel, Data Exfiltration via Email):

OfficeActivity
| where Operation == "Send" and RecipientScope == "External"
| where Attachments has_any ("pdf", "docx", "xlsx", "csv", "zip")
| extend AttachmentSize = toint(AttachmentsSize)
| where AttachmentSize > 52428800  // 50MB
| summarize TotalSize=sum(AttachmentSize), Count=count() by UserId, bin(TimeGenerated, 1h)
| where Count > 10 or TotalSize > 500000000  // 500MB
| project TimeGenerated, UserId, TotalSize, Count

Use Case 5: Malware and Ransomware

Detection Logic:

SignalDetection RuleSeverity
Known malware hashFile hash matches known malware IOCCritical
Signature detectionEDR/AV detects known malware familyCritical
Behavioral indicatorsProcess injection, LOLBIN execution (e.g., certutil downloading file)High
Ransomware file activityMass file extension changes (.encrypted, .locked) in short time windowCritical
Shadow copy deletionvssadmin delete shadows /all /quiet (Sysmon Event ID 1)Critical
Backup service stopService stop for backup or AV servicesHigh
Ransom note creationFile named "README.txt", "HOW_TO_DECRYPT.html" created en masseCritical
Network beaconingRegular outbound connections to same IP at fixed intervalsHigh

Use Case 6: DDoS Attack

Detection Logic:

SignalDetection RuleSeverity
Volume anomalyInbound requests > 10x baselineHigh
Source diversityRequests from > 1,000 unique source IPs in 1 minuteMedium
Slowloris / Slow POSTConnections with very slow request completionMedium
Application-layer floodHigh rate of login attempts or API calls from single sourceHigh
Reflected amplificationUDP traffic spike (DNS, NTP, SSDP) from spoofed sourcesCritical

Use Case 7: Insider Threat

Detection Logic:

SignalDetection RuleSeverity
Bulk file accessUser accesses > 500 files in a folder they normally don't accessHigh
Resignation indicatorUser accesses performance review, HR folder, or prints resume after hoursMedium
Data export before departureHigh data access in 30 days before resignationHigh
Privilege abuseAdmin accesses sensitive data they don't need for job functionHigh
Off-hours database accessDatabase login at 2 AM by non-DBA userHigh
External communicationEmail to personal address with attachments containing customer dataCritical
Print volume spikePrint volume > 3x personal baseline in a weekMedium

Use Case 8: Brute Force and Credential Stuffing

Detection Logic:

SignalDetection RuleSeverity
Multiple failed logins> 10 failed logins from single IP in 5 minutesHigh
Distributed brute forceSame username with failed logins from multiple IPsHigh
Credential stuffingSame IP with failed logins for multiple usernamesHigh
Successful login after brute forceSuccessful login after 10+ failures from same IPCritical
Password spraySingle password attempted across many accounts (low-and-slow)High
MFA fatigue attackMultiple MFA push notifications to same user in short windowHigh

Use Case 9: API Abuse

Detection Logic:

SignalDetection RuleSeverity
Rate limit violationsAPI client exceeding rate limits repeatedlyMedium
Authentication abuseMultiple API key failures from single clientHigh
Data scrapingSequential ID enumeration (e.g., /api/users/1, /api/users/2...)High
Unusual endpoint accessAPI endpoint accessed that this client has never usedMedium
Large payload requestsPOST/PUT requests with unusually large body sizesMedium
Off-hours API volumeAPI traffic > 5x baseline outside business hoursHigh
Error rate spike4xx/5xx error rate > 20% for an API endpointMedium

Use Case 10: Cloud Misconfiguration Exploitation

Detection Logic:

SignalDetection RuleSeverity
Public bucket createdS3 bucket ACL or policy changed to public (CloudTrail)Critical
Security group wide openInbound 0.0.0.0/0 added to security group (CloudTrail)Critical
IAM key creationNew IAM access key created for root or admin userHigh
Database public accessRDS instance modified to allow public accessCritical
Privilege escalation via IAMiam:AttachUserPolicy with AdministratorAccess policyCritical
CloudTrail disabledCloudTrail stopLogging or deleteTrail eventCritical
KMS key deletionkms:ScheduleKeyDeletion eventHigh
Unusual region usageAPI calls in region where organization has no presenceMedium

Detection Engineering

Detection engineering is the discipline of building, testing, and maintaining detection logic. It is the core technical function of a modern security monitoring program.

The Detection Engineering Lifecycle

┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐    ┌──────────────┐
│  RESEARCH    │ →  │  DEVELOP     │ →  │  TEST        │ →  │  DEPLOY      │ →  │  MAINTAIN    │
│  (Threat     │    │  (Write rule,│    │  (Unit test, │    │  (CI/CD,     │    │  (Tune, update│
│  intel,      │    │  query,      │    │  integration │    │  version     │    │  for new     │
│  incident    │    │  playbook)   │    │  test,       │    │  control,    │    │  variants,   │
│  post-mortem)│    │              │    │  red team)   │    │  rollback)   │    │  deprecate)  │
└──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘    └──────────────┘

Rule Writing Frameworks

Sigma

Sigma is the generic signature format for SIEM rules. It allows you to write one rule and convert it to Splunk SPL, KQL, Elastic DSL, QRadar AQL, and others.

Why Sigma matters:

  • Vendor-agnostic rule format
  • Large community repository (github.com/SigmaHQ/sigma)
  • Enables rule portability between SIEMs
  • Supports CI/CD for detection engineering

Example Sigma Rule (Suspicious PowerShell Download):

title: Suspicious PowerShell Download
logsource:
  product: windows
  service: sysmon
detection:
  selection:
    EventID: 1
    Image|endswith: '\powershell.exe'
    CommandLine|contains:
      - 'IEX(New-Object Net.WebClient)'
      - 'Invoke-Expression'
      - 'bitsadmin'
      - 'certutil -urlcache'
  condition: selection
falsepositives:
  - Legitimate administrative scripts
level: high

Sigma Rule Conversion:

## Convert to Splunk SPL
sigmac -t splunk rule.yml

## Convert to Microsoft Sentinel KQL
sigmac -t azure Sentinel rule.yml

## Convert to Elastic DSL
sigmac -t es-qs rule.yml

## Convert to QRadar AQL
sigmac -t qradar rule.yml

YARA

YARA is the pattern matching tool for malware identification and classification. While primarily for files, YARA-L is used in Chronicle for log-based detection.

YARA for File Detection:

rule APT_MALWARE_Example {
    meta:
        description = "Detects known APT malware family"
        author = "Security Team"
        date = "2026-06-15"
    strings:
    condition:
}

YARA-L for Log Detection (Chronicle):

rule suspicious_dns_tunneling {
  meta:
    author = "Security Team"
    description = "Detects potential DNS tunneling"
  events:
    
    
    condition:
}

Snort / Suricata Rules (Network Detection)

For organizations running IDS/IPS or network detection:

alert tcp any any -> any 443 (msg:"SUSPICIOUS TLS SNI - DGA Pattern"; 
    tls.sni; pcre:"/[a-z]{20,30}\.(com|net|org)/"; 
    sid:1000001; rev:1;)

Detection-as-Code

Detection-as-Code is the practice of managing detection rules with the same rigor as application code.

Principles:

  1. Version control: All rules in Git
  2. Peer review: Pull requests for rule changes
  3. Automated testing: Unit tests for every rule
  4. CI/CD deployment: Rules deployed via pipeline, not manual UI configuration
  5. Documentation: Every rule has a README explaining logic, false positives, and response
  6. Metrics: Track rule performance, false positive rate, and detection latency

Sample Detection-as-Code Repository Structure:

detections/
├── windows/
│   ├── privilege_escalation/
│   │   ├── local_admin_addition.yml
│   │   ├── local_admin_addition_test.py
│   │   └── README.md
│   ├── lateral_movement/
│   │   ├── psexec_execution.yml
│   │   └── wmi_lateral_movement.yml
│   └── malware/
│       ├── ransomware_behavior.yml
│       └── cobalt_strike_beacon.yml
├── linux/
│   ├── privilege_escalation/
│   ├── persistence/
│   └── lateral_movement/
├── cloud/
│   ├── aws/
│   ├── azure/
│   └── gcp/
├── network/
│   ├── dns_tunneling.yml
│   └── c2_beaconing.yml
└── tests/
    ├── data/
    │   ├── windows_security.log
    │   ├── cloudtrail.json
    │   └── sysmon.json
    └── test_runner.py

CI/CD Pipeline for Detection Rules:

## .github/workflows/detection-deploy.yml
name: Detection Deployment
on:
  push:
    branches: [main]
  pull_request:
    branches: [main]

jobs:
  test:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Install Sigma CLI
        run: pip install sigmatools
      - name: Validate Sigma Rules
        run: sigma check detections/
      - name: Convert and Test Rules
        run: |
          python tests/test_runner.py --target-splunk
          python tests/test_runner.py --target-sentinel
  
  deploy-splunk:
    needs: test
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Deploy to Splunk
        run: |
          python scripts/deploy_splunk.py --env production
        env:

  deploy-sentinel:
    needs: test
    if: github.ref == 'refs/heads/main'
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - name: Deploy to Sentinel
        run: |
          python scripts/deploy_sentinel.py --env production
        env:

Testing Detections

Every detection rule must be tested before deployment. Testing methods:

Test TypeHowPurposeFrequency
Unit testFeed known log samples to ruleValidate rule logicEvery rule change
Integration testTest in staging SIEM with real dataValidate in environmentWeekly
Red team testSimulate attack, verify detection firesValidate end-to-endQuarterly
Purple team testCollaborate with red team, validate and tuneTune and improveMonthly
False positive reviewReview alert output for benign triggersMaintain qualityWeekly

Red Team Test Example (Account Compromise):

## Step 1: Create test account
create_test_user "detection_test_user"

## Step 2: Simulate impossible travel
login_from "detection_test_user" "New York"    # 9:00 AM EST
login_from "detection_test_user" "Tokyo"     # 9:30 AM EST (impossible)

## Step 3: Verify detection fires within 15 minutes
assert_detection_fired "impossible_travel" "detection_test_user" within 15m

## Step 4: Cleanup
delete_test_user "detection_test_user"

False Positive Tuning

False positives destroy SOC credibility. The tuning process:

  1. Measure baseline: Run rule in "monitoring mode" for 2 weeks. Count alerts per day.
  2. Categorize: Tag each alert as true positive, false positive, or benign true positive.
  3. Identify patterns: Are false positives from specific systems, users, or time windows?
  4. Add exclusions: Add filter conditions to exclude known benign patterns.
  5. Re-test: Run for another 2 weeks. Target < 5% false positive rate.
  6. Document: Maintain exclusion documentation with business justification and expiry date.

Example Tuning Evolution:

## Version 1: Initial rule (too broad)
title: Off-Hours Login
detection:
  selection:
    EventID: 4624
    Hour: [0,1,2,3,4,5,22,23]
  condition: selection
## Result: 50 alerts/day. 90% false positive (night shift DBAs, offshore teams)

## Version 2: Add baseline exclusion
title: Off-Hours Login
detection:
  selection:
    EventID: 4624
    Hour: [0,1,2,3,4,5,22,23]
  filter:
    UserName|endswith: '_batch'  # Exclude service accounts
    UserName|startswith: 'dba_'   # Exclude known DBA accounts
  condition: selection and not filter
## Result: 10 alerts/day. 60% false positive (still catching offshore dev team)

## Version 3: Add geolocation and learning
title: Off-Hours Login
detection:
  selection:
    EventID: 4624
    Hour: [0,1,2,3,4,5,22,23]
  filter_known_users:
    UserName|endswith: '_batch'
  filter_baseline:
    UserName in known_offhours_users  # Dynamic lookup from UEBA baseline
  condition: selection and not filter_known_users and not filter_baseline
## Result: 3 alerts/day. 10% false positive. Actionable.

Coverage Mapping to MITRE ATT&CK

Map your detection rules to MITRE ATT&CK to identify gaps.

TacticTechniqueTechnique IDDetection RuleCoverage Status
Initial AccessPhishingT1566Email rule creation, suspicious attachment🟢 Covered
Initial AccessValid AccountsT1078Impossible travel, off-hours login🟢 Covered
ExecutionCommand and Scripting InterpreterT1059PowerShell logging, LOLBIN execution🟢 Covered
PersistenceAccount ManipulationT1098Admin group addition, email forwarding🟢 Covered
PersistenceCreate AccountT1136New local/domain account creation🟢 Covered
Privilege EscalationAbuse Elevation Control MechanismT1548UAC bypass, sudo abuse🟡 Partial
Defense EvasionIndicator RemovalT1070Log clearing (Event ID 1102)🟢 Covered
Credential AccessOS Credential DumpingT1003LSASS access, SAM dump🟢 Covered
DiscoveryAccount DiscoveryT1087AD enumeration, LDAP queries🔴 Not Covered
Lateral MovementRemote ServicesT1021RDP, SMB, PSExec, WMI🟢 Covered
CollectionData from Local SystemT1005Bulk file access, USB usage🟡 Partial
ExfiltrationExfiltration Over C2 ChannelT1041Beaconing detection, DNS tunneling🟢 Covered
ImpactData Encrypted for ImpactT1486Ransomware file extensions🟢 Covered

Gap Analysis Action:

  • If > 20% of techniques in a tactic are red, prioritize building rules for that tactic
  • Focus on techniques with high prevalence in your threat intelligence
  • Update coverage map quarterly

UEBA & Behavioral Analytics

User and Entity Behavior Analytics (UEBA) detects anomalies by establishing baselines of normal behavior and flagging deviations. It is essential for detecting insider threats, compromised credentials, and slow-and-low attacks that signature-based rules miss.

What UEBA Actually Does

FunctionDescriptionExample
Baseline establishmentLearns normal patterns for users, devices, and systems over time"User Alice normally logs in from London 9 AM–5 PM, accesses 20 files/day, sends 50 emails/day"
Anomaly detectionFlags deviations from baseline"Alice logged in from Moscow at 2 AM, accessed 500 files, and sent 0 emails"
Peer group analysisCompares user to similar users"Alice is in Finance. No one in Finance has ever accessed the Engineering code repository."
Risk scoringCombines multiple anomalies into a single risk score"Alice's risk score = 87/100 (high) due to impossible travel + off-hours access + bulk file access"
Entity trackingTracks risk across entities (user, device, IP, file)"IP 203.0.113.45 has risk score 95 due to 5 failed accounts + 2 successful + malware download"
Session stitchingLinks activities across sessions and devices"Alice's session on laptop + RDP session on server + AWS console login are the same attacker"

UEBA vs. Signature-Based Detection

AspectSignature-Based (SIEM Rules)UEBA (Behavioral)
What it detectsKnown bad patternsUnknown / unusual patterns
RequiresThreat intelligence, known IOCsBaseline data, ML models
False positive rateLow (if tuned) for known threatsHigher initially, improves with tuning
Time to valueImmediate (rule fires on match)2–4 weeks (needs baseline learning)
Best forKnown malware, brute force, known C2Insider threat, compromised credentials, APT
overheadLower (part of SIEM)Higher (dedicated UEBA tool or SIEM module)

Key UEBA Capabilities by Tool

ToolUEBA ApproachUnique StrengthWeakness
Splunk UBADedicated ML models, entity-centricExcellent entity resolution, risk scoringpremium-tier, requires separate license
Microsoft Sentinel UEBABuilt-in, free with Azure AD P2Native integration with Azure AD, Microsoft 365Limited to Microsoft ecosystem depth
ExabeamPure-play UEBA, timeline-basedStrong session stitching and timelinepremium-tier, complex deployment
SecuronixCloud-native UEBA + SIEMStrong cloud and SaaS coverageNewer, smaller ecosystem
IBM QRadar UBAIntegrated with QRadar SIEMFree with QRadar, good for insider threatLess advanced than Exabeam or Splunk UBA
Elastic SecurityML-based anomaly detectionOpen, customizable, part of Elastic stackRequires ML expertise to tune effectively
DarktraceSelf-learning AI (network + endpoint)Autonomous response, no rule writingBlack box, premium-tier, can be overly aggressive
Vectra AINetwork-focused AI detectionExcellent network detection, attacker lifecycle trackingNetwork-only, requires TAP/SPAN

Baseline Establishment Process

Phase 1: Data Collection (Week 1–2)

  • Collect 30 days of activity data: logins, file access, email, web, application usage
  • Ensure data quality: deduplicated, timestamped, actor-identified
  • Include metadata: time, location, device, application, data sensitivity

Phase 2: Baseline Learning (Week 3–4)

  • Calculate per-user baselines:
    • Login times (hours, days of week)
    • Geographic locations (cities, countries)
    • Devices used (laptop, mobile, VDI)
    • Applications accessed
    • File access volume and patterns
    • Data sensitivity levels accessed
    • Email volume and external recipients
    • Web categories visited
  • Calculate peer group baselines (department, role, geography)
  • Identify system-wide baselines (server traffic patterns, API baseline)

Phase 3: Threshold Setting (Week 5)

  • Set anomaly thresholds:
    • Statistical: 3-sigma deviation from mean
    • Peer-based: Deviation from peer group median
    • Temporal: Off-hours defined by individual work pattern
    • Contextual: Unusual combination of activities
  • Tune for false positives: Adjust thresholds until < 5% FP rate

Phase 4: Continuous Learning (Ongoing)

  • Re-baseline quarterly
  • Incorporate feedback from analyst investigations
  • Adjust for role changes, seasonal patterns (quarter-end for finance)
  • Retire old baselines for departed users

Risk Scoring Models

A practical risk scoring model combines multiple factors:

Risk Score = (Anomaly Score × Weight) + (Threat Intel Score × Weight) + (Asset Value × Weight)

Where:
- Anomaly Score (0–100): Deviation from baseline (login time, location, volume, peer group)
- Threat Intel Score (0–100): IOC match, known malicious IP, suspicious domain
- Asset Value (0–100): Sensitivity of data accessed (Public=0, Internal=20, Confidential=60, Highly Confidential=100)
- Weights: Typically 0.4, 0.3, 0.3

Example Risk Score Calculation:

FactorValueWeightContribution
Anomaly: Login from new country (Russia)800.432
Anomaly: Bulk file access (500 files vs. baseline 20)700.428
Threat Intel: IP is Tor exit node900.327
Asset Value: Accessed HR compensation data1000.330
Total Risk Score117 → Capped at 100

Risk Score Thresholds:

ScorePrioritySLAResponse
90–100Critical15 minutesImmediate SOC escalation, disable account pending review
70–89High1 hourSOC analyst investigation, notify manager
50–69Medium4 hoursQueue for daily review, automated email to user
30–49Low24 hoursWeekly trend review, no immediate action
< 30InformationalNoneFeed into baseline improvement

Insider Threat Detection with UEBA

Insider threats are the hardest to detect because the actor has legitimate access. UEBA excels here.

Insider Threat TypeBehavioral IndicatorsUEBA Detection
The ResignerAccessing HR folders, printing resume, off-hours access to IPBaseline deviation: printing volume, unusual file access categories
The Data ThiefBulk download of customer data, USB usage, email to personal addressVolume anomaly + data sensitivity + exfiltration channel
The SaboteurAdmin access to critical systems, configuration changes, backup deletionPrivileged action anomaly + off-hours + system criticality
The Negligent InsiderSharing password, falling for phishing, using unapproved cloud storagePolicy violation + credential exposure indicators
The Compromised InsiderAccount used by attacker from new location, simultaneous loginImpossible travel + credential abuse pattern

Privacy Considerations:

  • UEBA monitoring must be documented in privacy policy and employee agreements
  • In EU, works councils may need to be consulted
  • Avoid keystroke logging or screen recording without explicit legal basis
  • Focus on data access patterns, not content (e.g., "accessed 500 files" not "read confidential merger document")
  • Implement data minimization: only collect data needed for security
  • Set retention limits: behavioral baselines should be deleted when user departs

💡 Tip from Singahi: UEBA is not a "set and forget" tool. It requires 2–4 weeks of learning, then continuous tuning. Organizations that buy Exabeam or Splunk UBA and never tune it end up with 200 daily alerts and an ignored queue. Assign a dedicated analyst for the first 90 days to provide feedback and tune thresholds.


Cloud Monitoring

Cloud monitoring is no longer optional. For most organizations, cloud infrastructure is the primary, or only, infrastructure. Monitoring cloud-native services requires cloud-native tools and techniques.

AWS CloudTrail

CloudTrail is the audit log of AWS. Every API call is recorded.

Critical CloudTrail Configurations:

SettingRecommendationWhy
Multi-region trailEnable in all regionsAttackers often operate in unused regions to evade detection
Organization trailEnable at organization levelCaptures all accounts in AWS Organization without per-account setup
Data eventsEnable for S3 and LambdaRead/write object-level events and Lambda function invocation
Insight eventsEnableAutomatically detects unusual API call volumes (e.g., 20x baseline)
Log file validationEnableCryptographically proves logs haven't been tampered
S3 bucket securitySeparate account, versioning, MFA delete, access loggingProtects the audit log from deletion by compromised admin
EncryptionSSE-KMS with customer-managed keyPrevents AWS from accessing your logs

CloudTrail Events to Alert On Immediately:

Event NameRiskDetection Priority
ConsoleLogin with MFAUsed: false and Root accountRoot account without MFACritical
CreateAccessKey for root or admin userNew API keys for privileged usersCritical
DeleteTrail or StopLoggingCloudTrail tamperingCritical
PutBucketPolicy / PutBucketAcl with public accessPublic S3 bucketCritical
AuthorizeSecurityGroupIngress with 0.0.0.0/0Open security groupCritical
AttachUserPolicy with AdministratorAccessPrivilege escalationHigh
CreateUser / CreateRole in unused regionStealth account creationHigh
DeleteBucket / DeleteBucketPolicyData destructionHigh
PutBucketVersioning with Status: SuspendedDisabling versioning for ransomwareHigh
PutBucketLifecycle with short expirationData destruction via lifecycleMedium

CloudTrail Log Analysis Example (Athena):

SELECT 
    eventTime,
    eventName,
    userIdentity.arn,
    sourceIPAddress,
    requestParameters
FROM cloudtrail_logs
WHERE eventName IN ('DeleteTrail', 'StopLogging', 'PutBucketPolicy')
    AND eventTime > current_timestamp - interval '7' day
ORDER BY eventTime DESC;

AWS GuardDuty

GuardDuty is AWS's managed threat detection service. It uses ML and threat intelligence.

Finding TypeDescriptionAction
CredentialAccess:IAMUser/AnomalousBehaviorUnusual IAM user API activityInvestigate user, check for compromise
CryptoCurrency:EC2/BitcoinTool.B!DNSEC2 instance communicating with Bitcoin domainLikely compromised instance, isolate
DefenseEvasion:IAMUser/CloudTrailLoggingDeletedCloudTrail deletedCritical incident, investigate immediately
Discovery:S3/AnomalousBehaviorUnusual S3 API activityCheck for data reconnaissance
Exfiltration:IAMUser/AnomalousBehaviorUnusual data transfer patternsInvestigate for data exfiltration
Impact:EC2/MaliciousDomainRequest.ReputationEC2 requesting known malicious domainCompromised instance
PenTest:IAMUser/KaliLinuxAPI calls from Kali LinuxCould be pen test or attacker using Kali
Persistence:IAMUser/AnomalousBehaviorIAM user persistence behaviorCheck for new users, roles, policies
Recon:IAMUser/NetworkPermissionsIAM user checking network permissionsReconnaissance for lateral movement
Stealth:S3/ServerAccessLoggingDeletedS3 logging disabledCheck for cover-up activity

GuardDuty Integration:

  • Forward findings to SIEM via EventBridge → Kinesis → Lambda → SIEM
  • Use Security Hub for centralized AWS security findings
  • Integrate with SOAR for automated response (e.g., isolate EC2 on GuardDuty finding)

Azure Monitor & Sentinel

Azure's monitoring stack is deeply integrated. Azure Monitor collects logs and metrics. Sentinel analyzes them.

Azure Monitor Data Sources:

SourceData TypeCollection Method
Azure Activity LogControl plane events (resource creation, RBAC)Native, automatically collected
Azure AD Sign-in LogsAuthentication eventsAzure AD Diagnostic Settings → Log Analytics
Azure AD Audit LogsDirectory changes (user, group, app)Azure AD Diagnostic Settings → Log Analytics
Azure VMOS events, performanceAzure Monitor Agent or Log Analytics agent
Azure NSGNetwork flow logsNSG Flow Logs → Storage Account → Log Analytics
Azure FirewallTraffic logs, application rulesDiagnostic Settings → Log Analytics
Azure Key VaultAccess and audit eventsDiagnostic Settings → Log Analytics
Azure StorageStorage analytics logsStorage diagnostic settings
Azure KubernetesControl plane, container logsContainer Insights, Diagnostic Settings
Microsoft 365Exchange, SharePoint, Teams auditMicrosoft 365 Defender connector
Non-Azure sourcesOn-prem, AWS, GCPLog Analytics agent, CEF collector, AMA, Azure Arc

Azure Sentinel Key Tables:

TableContentUse Case
SecurityAlertAlerts from Microsoft and partner productsCentralized alert management
SecurityEventWindows Security events (Sysmon, 4688, etc.)Endpoint detection
SigninLogsAzure AD sign-in eventsAccount compromise detection
AuditLogsAzure AD directory changesPrivilege escalation
AzureActivityAzure control planeCloud misconfiguration exploitation
CommonSecurityLogCEF-formatted logs (firewalls, proxies)Network security
DeviceLogonEventsDefender for Endpoint logon eventsEndpoint detection
DeviceNetworkEventsDefender for Endpoint network eventsNetwork beaconing, C2
DeviceFileEventsDefender for Endpoint file eventsMalware, ransomware
OfficeActivityMicrosoft 365 auditInsider threat, email exfiltration
AWSCloudTrailAWS CloudTrail via connectorMulti-cloud monitoring

GCP Security Command Center

GCP's security monitoring is centered around Security Command Center (SCC).

SCC FeatureDescriptionTier
Asset InventoryContinuous discovery of GCP resourcesFree
Security SourcesFindings from Security Health Analytics, Web Security Scanner, Container AnalysisFree
Vulnerability DiscoveryOS and container vulnerability scanningPremium
Threat DetectionBuilt-in threat detection (Crypto mining, brute force, data exfiltration)Premium
Event Threat DetectionDetects threats in Cloud Logging (analogous to GuardDuty)Premium
Container Threat DetectionDetects container runtime threats (analogous to Falco)Premium

GCP Cloud Audit Logs (Critical for Monitoring):

Log TypeScopeKey Events
Admin ActivityAdministrative actions on resourcesIAM changes, resource creation/deletion, policy changes
Data AccessRead/write operations on user dataBigQuery queries, Cloud Storage object access, SQL queries
System EventAutomated Google system eventsInstance migration, maintenance, encryption key rotation
Policy DeniedAccess denied due to IAM policyFailed authorization attempts, reconnaissance

GCP Log Analysis Example (BigQuery):

SELECT
  timestamp,
  protoPayload.authenticationInfo.principalEmail,
  protoPayload.methodName,
  protoPayload.resourceLabels.project_id,
  protoPayload.serviceData.policyDelta.auditConfigDeltas
FROM `project-id.audit_logs.cloudaudit_googleapis_com_data_access`
WHERE protoPayload.methodName LIKE 'setIamPolicy'
  AND timestamp > TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL 7 DAY)
ORDER BY timestamp DESC;

Multi-Cloud Monitoring Strategy

Most mid-to-large organizations use multiple clouds. A unified monitoring approach is essential.

Architecture Pattern: Centralized Cloud SIEM with Cloud-Native Feeds

┌────────────────────────────────────────────────────────────────────┐
│                    CENTRAL SIEM (Splunk / Sentinel / Elastic)       │
│  ┌─────────────┐  ┌─────────────┐  ┌─────────────┐  ┌──────────┐ │
│  │ Correlation │  │ Detection   │  │ Cases       │  │ Reporting│ │
│  │ Engine      │  │ Rules       │  │             │  │          │
│  └─────────────┘  └─────────────┘  └─────────────┘  └──────────┘ │
└────────────────────────────────────────────────────────────────────┘
                              ▲
        ┌─────────────────────┼─────────────────────┐
        │                     │                     │
┌───────┴───────┐   ┌─────────┴─────────┐   ┌───────┴───────┐
│   AWS LANDING │   │   AZURE LANDING   │   │   GCP LANDING │
│   ZONE        │   │   ZONE            │   │   ZONE        │
│               │   │                   │   │               │
│ CloudTrail →  │   │ Activity Log →    │   │ Cloud Audit → │
│ Kinesis →    │   │ Event Hub →        │   │ Pub/Sub →     │
│ Lambda →     │   │ Function →          │   │ Function →    │
│ SIEM         │   │ SIEM                │   │ SIEM          │
│               │   │                   │   │               │
│ GuardDuty →  │   │ Defender for →      │   │ SCC →         │
│ SecurityHub →│   │ Cloud → SIEM        │   │ SIEM          │
│ SIEM         │   │                   │   │               │
└───────────────┘   └───────────────────┘   └───────────────┘

Multi-Cloud Monitoring Challenges:

ChallengeSolution
Different log formatsNormalize to common schema (ECS or ASIM) at ingestion
Different event modelsCreate unified detection rules that work across cloud event types
Cross-cloud correlationUse SIEM correlation engine with entity resolution (same user across AWS + Azure)
Data egress overheadUse cloud-native aggregation before cross-cloud transfer; use private interconnect
IAM for monitoringCreate least-privilege monitoring roles in each cloud with read-only access
Alert fatigueCentralize alert management, deduplicate cross-cloud findings

CSPM Integration with Monitoring

Cloud Security Posture Management (CSPM) finds misconfigurations. Monitoring detects exploitation. They must work together.

CSPM ToolMonitoring IntegrationValue
Prisma CloudSends alerts to SIEM, correlates with runtime threatsDetects misconfiguration → exploitation chain
WizIntegrates with SIEM, provides graph-based riskExploit path visualization
Orca SecuritySide-scanning + alert integrationAgentless monitoring with threat detection
Microsoft Defender for CloudNative Sentinel integrationUnified Azure security posture + threat detection
AWS Security HubCentralizes GuardDuty, Inspector, Macie findingsSingle pane for AWS security
Datadog CSPMUnified with Datadog Security MonitoringSecurity + observability in one platform

Example CSPM → Monitoring Workflow:

  1. CSPM detects S3 bucket is public (prisma:public-s3-bucket)
  2. SIEM receives CSPM alert, creates low-priority ticket
  3. CloudTrail shows GetObject operations on that bucket from unknown IP
  4. SIEM correlates CSPM alert + CloudTrail anomaly → Critical alert
  5. SOAR automatically blocks IP and restricts bucket policy

💡 Tip from Singahi: The biggest cloud monitoring gap we see is organizations enabling CloudTrail but never reviewing it. CloudTrail without alerting is just premium-tier storage. At minimum, set up alerts for ConsoleLogin, DeleteTrail, PutBucketPolicy, and AuthorizeSecurityGroupIngress. These 4 events cover 80% of critical cloud security incidents.


Kubernetes & Container Monitoring

Containers and Kubernetes have unique monitoring challenges: ephemeral infrastructure, shared kernels, and complex networking. Standard host-based monitoring is insufficient.

Kubernetes Audit Logs

K8s audit logs record every request to the API server. They are the CloudTrail of Kubernetes.

Enabling K8s Audit Logs:

For managed K8s (EKS, AKS, GKE), audit logs are partially enabled. For self-managed K8s, you must configure the API server.

## /etc/kubernetes/audit-policy.yaml
apiVersion: audit.k8s.io/v1
kind: Policy
rules:
  # Log all requests at the Metadata level
  - level: Metadata
    omitStages:
      - RequestReceived
  
  # Log pod changes at RequestResponse level
  - level: RequestResponse
    resources:
      - group: ""
        resources: ["pods", "pods/status", "pods/log"]
    omitStages:
      - RequestReceived
  
  # Log secret access at RequestResponse
  - level: RequestResponse
    resources:
      - group: ""
        resources: ["secrets"]
    omitStages:
      - RequestReceived
  
  # Log RBAC changes at RequestResponse
  - level: RequestResponse
    resources:
      - group: "rbac.authorization.k8s.io"
        resources: ["roles", "rolebindings", "clusterroles", "clusterrolebindings"]
    omitStages:
      - RequestReceived
  
  # Log exec and attach at RequestResponse
  - level: RequestResponse
    resources:
      - group: ""
        resources: ["pods/exec", "pods/attach"]
    omitStages:
      - RequestReceived
  
  # Don't log health checks
  - level: None
    nonResourceURLs:
      - /healthz
      - /livez
      - /readyz

Critical K8s Audit Events to Monitor:

VerbResourceRiskDetection Priority
createpodsNew pod creation (could be malicious workload)Medium
deletepodsPod deletion (disruption)Medium
createpods/execexec into running pod (lateral movement, data access)High
createpods/attachattach to running podHigh
createsecretsSecret creation (could be exfiltration staging)Medium
getsecretsSecret reading (credential theft)High
createroles / rolebindingsPrivilege escalation within namespaceHigh
createclusterroles / clusterrolebindingsCluster-wide privilege escalationCritical
createserviceaccountsNew service account (persistence)Medium
patchdeploymentsDeployment modification (could inject sidecar)High
deleteeventsEvent deletion (cover-up)Critical
deleteauditpoliciesAudit policy tamperingCritical

Example K8s Audit Alert (Falco + Fluent Bit):

## falco rule
- rule: Unauthorized K8s Secret Access
  desc: Detect reading secrets by non-service-account users
  condition: >
    k8s.audit.user != "system:serviceaccount:*" and
    k8s.audit.verb in ("get", "list") and
    k8s.audit.objectref.resource = "secrets"
  output: "Unauthorized secret access by user=%k8s.audit.user resource=%k8s.audit.objectref.name"
  priority: HIGH

Falco, Runtime Security

Falco is the open-source runtime security engine for containers. It detects unexpected behavior using system calls.

Falco Deployment Options:

MethodProsCons
DaemonSetStandard K8s deployment, easy to manageResource overhead on every node
SidecarPer-pod monitoring, no node access neededMore complex, higher total resource
eBPF probeModern, no kernel moduleRequires recent kernel, eBPF support
Kernel moduleTraditional, broad compatibilityKernel compatibility issues, less secure
gVisorSandbox-level monitoringPerformance overhead, limited adoption

Critical Falco Rules for Production:

## falco_rules.local.yaml
- rule: Privileged Container Started
  desc: Detect privileged container creation
  condition: >
    spawned_process and
    container.privileged = true
  output: "Privileged container started: %container.name"
  priority: CRITICAL

- rule: Sensitive File Access in Container
  desc: Detect access to /etc/shadow, /etc/passwd
  condition: >
    spawned_process and
    (proc.name in ("cat", "less", "more", "vim")) and
    (fd.name contains "/etc/shadow" or fd.name contains "/etc/passwd")
  output: "Sensitive file access: %proc.name %fd.name %container.name"
  priority: HIGH

- rule: Outbound Connection from Sensitive Container
  desc: Database containers should not make outbound connections
  condition: >
    outbound and
    container.name contains "database" and
    fd.type = "ipv4"
  output: "Outbound connection from database container: %container.name %fd.name"
  priority: HIGH

- rule: Shell Spawned in Container
  desc: Detect interactive shell in container
  condition: >
    spawned_process and
    shell_procs and
    container.id != "host"
  output: "Shell spawned in container: %proc.name %container.name"
  priority: MEDIUM

Admission Controllers

Admission controllers intercept requests to the K8s API server before persistence. They enforce policy.

ControllerPurposeSecurity Value
OPA GatekeeperPolicy enforcement using Rego languagePrevent privileged containers, enforce resource limits, require labels
KyvernoNative K8s policy engineEasier to write than Rego, mutate resources, validate configs
Pod Security Admission (PSA)Built-in K8s pod securityEnforce Pod Security Standards (Privileged, Baseline, Restricted)
Network PoliciesL3/L4 network segmentationPrevent lateral movement between pods and namespaces
Cert-ManagerAutomated TLS certificate managementEnsure all services use valid TLS

Kyverno Policy Example (Block Privileged Containers):

apiVersion: kyverno.io/v1
kind: ClusterPolicy
metadata:
  name: disallow-privileged
spec:
  validationFailureAction: enforce
  rules:
  - name: validate-privileged
    match:
      resources:
        kinds:
        - Pod
    validate:
      message: "Privileged containers are forbidden"
      pattern:
        spec:
          containers:
          - securityContext:
              =(privileged): "false"

Container Runtime Security

Beyond Falco, additional runtime security layers:

ToolApproachUse Case
SysdigSystem call capture + analysisDeep forensics, compliance recording
TrivyContainer image scanningDetect CVEs in images before deployment
Snyk ContainerImage + dependency scanningDeveloper-friendly vulnerability management
AnchorePolicy-based image scanningEnterprise image compliance
Aqua SecurityFull container security platformRuntime protection, image scanning, network micro-segmentation
Twistlock (Prisma Cloud)Palo Alto container securityIntegrated with Prisma Cloud for full cloud security

Service Mesh Monitoring

Service meshes (Istio, Linkerd) provide mTLS, traffic management, and observability.

Mesh FeatureSecurity Monitoring Value
mTLS everywhereEvery service-to-service call is authenticated. Monitor for mTLS failures (policy violations).
Access policiesAuthorizationPolicy resources control who can talk to whom. Monitor for policy denials.
Request telemetryEvery request generates metrics and logs. Monitor for unusual latency, error rates, or request volumes.
Ingress gatewayCentralized entry point. Monitor for DDoS, injection attempts, unusual geographic distribution.

Istio Security Monitoring:

## Monitor for mTLS policy violations
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
  name: default
  namespace: production
spec:
  mtls:
    mode: STRICT
## Any connection without mTLS will be rejected and logged

Network Policies

K8s NetworkPolicies are firewall rules for pods. Monitoring them is essential.

Policy TypePurposeMonitor For
Deny-all defaultBlock all ingress/egress by defaultAny traffic that should be blocked but isn't
Namespace isolationOnly allow same-namespace trafficCross-namespace traffic (lateral movement)
Egress controlRestrict outbound connectionsUnexpected outbound connections (C2, exfiltration)
Ingress whitelistOnly allow specific sourcesConnections from unauthorized sources

Monitoring Network Policy Violations:

  • Calico: calico-policy-controller logs denied packets
  • Cilium: cilium monitor shows policy drops
  • Weave: weave-npc logs denied connections
  • Log these to SIEM and alert on denied connection attempts from sensitive namespaces

💡 Tip from Singahi: Most organizations deploying Kubernetes monitor pod CPU and memory but not K8s audit logs. The K8s API server is the control plane, every attack on K8s goes through it. Enable K8s audit logs, forward them to your SIEM, and alert on pods/exec, secrets/get, and clusterrolebindings/create. These 3 events catch 80% of K8s attacks.


Database Monitoring

Databases contain the crown jewels. Monitoring database activity is non-negotiable for ISO 27001 compliance.

Database Audit Log Categories

CategoryWhat to LogWhy
AuthenticationSuccessful and failed logins, logouts, session durationAccount compromise detection
AuthorizationPrivilege changes, role grants, permission changesPrivilege escalation detection
Schema changesCREATE, ALTER, DROP table, index, viewUnauthorized data structure changes
Data accessSELECT, INSERT, UPDATE, DELETE on sensitive tablesData theft, unauthorized access
Data changesDML on sensitive tablesData tampering, integrity violations
Admin activityDBA commands, backup/restore, configuration changesInsider threat, operational risk
Query performanceSlow queries, full table scans, connection spikesReconnaissance, DoS, abuse

Database Activity Monitoring (DAM) Tools

DAM tools sit between the database and the application, monitoring all traffic without modifying the database.

ToolDeploymentDatabase SupportKey Strengthlicensing Range
GreenSQLProxyMySQL, PostgreSQLOpen-source option, basic SQL injection detectionFree

Native Database Audit Configuration

PostgreSQL (pgaudit):

-- Install pgaudit extension
CREATE EXTENSION pgaudit;

-- Configure pgaudit in postgresql.conf
pgaudit.log = 'write, ddl, role'
pgaudit.log_catalog = off
pgaudit.log_parameter = on
pgaudit.log_statement_once = off
pgaudit.log_level = log

-- Audit specific tables
CREATE TABLE sensitive_customers (...);
ALTER TABLE sensitive_customers SET (pgaudit.log = 'all');

Microsoft SQL Server (SQL Server Audit):

-- Create server audit
USE master;
CREATE SERVER AUDIT ComplianceAudit
TO FILE (FILEPATH = 'D:\SQLAudit\', MAXSIZE = 100 MB, MAX_ROLLOVER_FILES = 10)
WITH (ON_FAILURE = CONTINUE);
ALTER SERVER AUDIT ComplianceAudit WITH (STATE = ON);

-- Create database audit specification
USE ProductionDB;
CREATE DATABASE AUDIT SPECIFICATION SensitiveDataAudit
FOR SERVER AUDIT ComplianceAudit
ADD (SELECT, INSERT, UPDATE, DELETE ON dbo.Customers BY public)
WITH (STATE = ON);

Oracle (Unified Audit Policy):

-- Create unified audit policy
CREATE AUDIT POLICY monitor_sensitive_data
ACTIONS SELECT, INSERT, UPDATE, DELETE ON hr.employees,
        SELECT, INSERT, UPDATE, DELETE ON hr.salary_data
WHEN 'SYS_CONTEXT(''USERENV'', ''SESSION_USER'') != ''HR_APP'''
EVALUATE PER STATEMENT;

-- Enable policy
AUDIT POLICY monitor_sensitive_data;

Query Analysis for Security

Monitor query patterns for signs of abuse:

PatternDetectionSecurity Concern
SELECT * FROM customers without WHEREFull table scan on sensitive tableData scraping, unauthorized bulk access
UNION SELECT in application querySQL injection attemptActive attack
INTO OUTFILE / BULK INSERTData export to fileData exfiltration preparation
DROP TABLE, TRUNCATE TABLEDestructive DDLSabotage, ransomware
CREATE USER, GRANT DBAPrivilege escalationUnauthorized admin access
Queries from unusual applicationApp not in whitelist connecting to DBCompromised application server
Connection spike from single IP>100 connections in 1 minuteBrute force, DoS
Long-running queries during off-hoursQuery running >1 hour at 2 AMBatch data theft

Database Access Pattern Monitoring

Establish baselines for database access:

Baseline ElementHow to MeasureAnomaly Detection
Connection sourcesWhich IPs/applications connectNew source IP = anomaly
Connection timesWhen is database normally accessedOff-hours access = anomaly
Query volumeQueries per hour per user/application3x volume spike = anomaly
Table access patternsWhich tables each app/user normally queriesNew table access = anomaly
Query complexityAverage rows returned, join countFull table scan on large table = anomaly
Privilege usageWhich privileges are normally usedElevated privilege use = anomaly

DLP for Databases

Database DLP prevents sensitive data from leaving the database environment.

DLP FunctionImplementationExample Tool
Data classificationTag columns with sensitivity levelImperva, IBM Guardium, Microsoft Purview
Query inspectionBlock or alert on queries returning sensitive dataImperva SecureSphere, Oracle Database Vault
Masking / redactionReturn masked data to unauthorized usersOracle Data Redaction, IBM Guardium, Dynamic Data Masking (SQL Server)
BlockingPrevent unauthorized SELECT on sensitive tablesDatabase firewall (Imperva, GreenSQL)
AlertingAlert when sensitive data is accessed in bulkDAM tools, native audit

💡 Tip from Singahi: Database monitoring is the most commonly neglected area in ISO 27001 implementations. Organizations monitor firewalls and endpoints but treat databases as "internal" and safe. The majority of data breaches involve database compromise. At minimum, enable native database audit logs, forward them to your SIEM, and alert on failed logins, privilege changes, and off-hours access. Native audit is free, use it.


Network Monitoring

Network monitoring is the foundation of security visibility. Even in a zero-trust world, network telemetry reveals attacker movement, data exfiltration, and command-and-control communication.

NetFlow, sFlow, and IPFIX

Flow protocols summarize network traffic metadata without capturing full packets.

ProtocolLayerKey FieldsBest ForTypical Collector
NetFlow v5L3/L4Source/dest IP, port, protocol, bytes, packetsBasic traffic analysis, Cisco environmentsElastiFlow, ntopng, Plixer
NetFlow v9L3/L4 + optionsFlexible templates, application IDs, QoSModern Cisco, detailed traffic categorizationElastiFlow, Cisco Stealthwatch
IPFIXL3/L4 + extensibleStandardized extensible format, deep packet inspection fieldsMulti-vendor, modern networksElastiFlow, ntopng, Flowmon
sFlowL2/L3/L4Packet sampling + counter samplingHigh-speed networks, real-time visibilitysFlow-RT, ntopng, InMon
VPC Flow LogsL3/L4Cloud-native version of NetFlowAWS, Azure, GCP network visibilityCloud-native (Athena, BigQuery) or SIEM

Flow Analysis for Security:

AnomalyFlow IndicatorAttack Phase
C2 beaconingRegular 60-second connections to same external IPPost-exploitation
Data exfiltrationLarge outbound flow to new external IPExfiltration
Lateral movementInternal flows between unusual host pairsLateral movement
Port scanMany destination ports from single sourceReconnaissance
DDoSMany source IPs, same destination, high packet rateImpact
TunnelingDNS or ICMP flows with high byte countExfiltration / C2
Tor exit nodeConnection to known Tor exit listAnonymization

Packet Capture (Full PCAP)

Full packet capture provides the deepest visibility but at the highest overhead.

Use CaseCapture StrategyRetentionTools
Incident responseCapture on-demand or triggered by alert24–72 hourstcpdump, Wireshark, Moloch/Arkime
Threat huntingSelective capture on critical segments7–30 daysArkime, Stenographer, Netresec
ComplianceFull capture on regulated segments30–90 daysEndace, Niksun, Gigamon
Performance analysisSampled capture (1:1000)1–7 daysntopng with nDPI

Arkime (Open-Source PCAP Analysis):

  • High-performance full packet capture and indexing
  • Web-based search and session reconstruction
  • Supports 10Gbps+ capture rates
  • Integrates with YARA for file extraction
  • overhead: free (open source), requires significant storage infrastructure

Storage Calculation for PCAP:

1 Gbps network = 125 MB/s = 450 GB/hour = 10.8 TB/day = 324 TB/month
At 10 Gbps: 108 TB/day = 3.2 PB/month

Solution: Capture only on critical segments, use flow for everything else, trigger PCAP on alert.

IDS / IPS

Intrusion Detection Systems (IDS) and Intrusion Prevention Systems (IPS) inspect traffic for known attack signatures.

ToolTypeDeploymentBest Foroverhead
SnortIDS/IPSNetwork sensor, inline or tapOpen-source IDS, community rulesFree (community rules)
SuricataIDS/IPS/NSMNetwork sensor, high performanceHigh-speed networks, full packet capture integrationFree (ET rules)
Zeek (Bro)NSMNetwork sensor, passiveProtocol analysis, behavioral detection, traffic parsingFree

Suricata Rule Example (C2 Beaconing):

    msg:"SUSPICIOUS TLS SNI - DGA Pattern";
    classtype:trojan-activity; sid:1000001; rev:1;
)

Zeek Script Example (Detect RDP over Non-Standard Port):

## rdp_anomaly.zeek
event RDP::client_connect(c: connection)
    {
        {
        }
    }

Network Detection and Response (NDR)

NDR tools use AI/ML to detect network anomalies without signatures or decryption.

ToolApproachDeploymentKey Strengthlicensing Range

NDR vs. IDS/IPS:

AspectIDS/IPSNDR
Detection methodSignatures, known patternsBehavioral, ML-based, anomaly detection
Encrypted trafficBlind to encrypted (unless SSL inspection)Can detect anomalies in encrypted traffic metadata
Known vs. unknownKnown threatsUnknown threats
DeploymentInline or passivePassive only (TAP/SPAN)
ResponseCan block (IPS)Alerts only (response via SOAR or firewall integration)
MaintenanceRequires rule updatesRequires tuning and learning period
overheadLowerHigher

Best practice: Run both. IDS/IPS for known threats and blocking. NDR for unknown threats and behavioral detection.

Encrypted Traffic Analysis

With 90%+ of web traffic encrypted, traditional deep packet inspection is increasingly ineffective.

TechniqueWhat It RevealsImplementation
TLS fingerprintingClient hello parameters, cipher suites, SNIJA3/JA3S hashes. Malware often has unique TLS fingerprints.
SNI analysisDomain name in TLS handshakeBlock or alert on suspicious SNI patterns (DGA, young domains)
Traffic timing analysisInter-arrival times, flow durationBeaconing detection without decrypting payload
Packet size analysisByte distribution, entropyEncrypted C2 often has distinct packet size patterns
DNS over HTTPS (DoH) detectionDoH resolver IPs, timing patternsBlock known DoH IPs, detect anomalous HTTPS to DNS-like ports
SSL/TLS inspectionFull decryption for inspectionLegal/privacy concerns, certificate pinning breaks, performance impact

JA3 Fingerprinting Example:

## JA3 is a TLS fingerprint based on Client Hello
## Normal Chrome: ja3=769,47-53-5-10-49161-49162-49171-49172-50-56-19-4,0-10-11,23-24-25,0
## Metasploit: ja3=769,53-47-5-10-49161-49162-49171-49172-50-56-19-4,65281-0-11-35-5-16,23-24-25,0
## Alert when JA3 matches known malware fingerprint

DNS Monitoring

DNS is the most underutilized monitoring source. 80% of malware uses DNS for C2.

DNS Monitoring TechniqueWhat It DetectsImplementation
Query loggingAll DNS queries and responsesEnable logging on DNS servers (BIND, Windows DNS, Infoblox)
NXDOMAIN rateDGA domains (malware generating random domains)Alert on >100 NXDOMAIN/hour from single host
TXT record sizeDNS tunneling (data encoded in TXT records)Alert on TXT records >200 bytes
Query volume anomalyDNS tunneling or C2 beaconingBaseline queries per host, alert on 10x spike
Domain ageNewly registered domainsAlert on queries to domains < 7 days old
Threat intelligenceKnown malicious domainsIntegrate DNS logs with threat intel feeds (e.g., MISP)
DNS over HTTPS / TLSBypassing DNS monitoringBlock DoH/DoT at firewall, detect anomalous HTTPS to port 443 of known DNS IPs

DNS Logging Configuration (Windows DNS):

## Enable debug logging (not for production scale, use analytic logs for scale)
dnscmd /Config /LogLevel 0x8100F331
## Or use Windows DNS Analytical Logs (Event ID 1000+)
## Enable in DNS Manager → Properties → Event Logging → Analytical

DNS Monitoring Alert Example (KQL):

// Detect potential DGA activity
DNS_Logs
| where TimeGenerated > ago(1h)
| summarize QueryCount=count(), UniqueDomains=dcount(QueryName) by ClientIP
| where QueryCount > 100 and UniqueDomains > 50
| extend DGAScore = QueryCount * UniqueDomains
| where DGAScore > 5000
| project ClientIP, QueryCount, UniqueDomains, DGAScore
| sort by DGAScore desc

💡 Tip from Singahi: If you do one thing for network monitoring this quarter, enable DNS query logging and forward it to your SIEM. DNS is free telemetry that catches malware C2, data exfiltration, and phishing. Most organizations already have DNS infrastructure, they just don't log it. BIND, Windows DNS, and Infoblox all support query logging. Turn it on.


Endpoint Monitoring

Endpoints are where attacks begin and where data resides. EDR/XDR provides the granular visibility needed to detect and respond to endpoint threats.

EDR / XDR Landscape

ProductVendorDeploymentKey Strengthlicensing Range
WazuhOpen sourceOpen source agentFree, includes FIM, rootkit detection, active responseFree

EDR Data Sources

Modern EDRs collect extensive telemetry:

Telemetry TypeWhat It CapturesSecurity Value
Process executionEvery process start, parent-child relationships, command lineMalware execution, LOLBIN usage, living-off-the-land
File system activityFile creation, modification, deletion, renamingRansomware file activity, data staging, persistence
Registry changesRegistry key creation, modification, deletionPersistence mechanisms, configuration changes
Network connectionsEvery outbound connection, DNS query, listening portC2 beaconing, data exfiltration, lateral movement
DLL/module loadsEvery DLL loaded into a processDLL injection, hooking, code injection
Memory injectionWriteProcessMemory, VirtualAllocEx, CreateRemoteThreadProcess hollowing, code injection
AuthenticationLogon events, token escalation, privilege changesPass-the-hash, token impersonation, privilege escalation
Script executionPowerShell, WMI, VBScript, JavaScript, PythonLiving-off-the-land, script-based malware
API callsKey Windows API calls (NtCreateThreadEx, etc.)Advanced malware behavior

Sysmon, The Essential Windows Telemetry

Sysmon is a free Microsoft tool that provides deep Windows telemetry. Every EDR uses Sysmon-like data. Even with a commercial EDR, Sysmon provides independent visibility.

Critical Sysmon Event IDs:

Event IDDescriptionDetection Value
1Process creationBaseline execution, LOLBIN detection, malware execution
2Process changed file creation timeTimestomping (anti-forensics)
3Network connectionC2 beaconing, data exfiltration, lateral movement
5Process terminatedProcess lifecycle, cleanup after infection
6Driver loadedRootkit loading, malicious driver
7Image loadedDLL injection, malicious DLL loading
8CreateRemoteThread detectedCode injection, process hollowing
9RawAccessRead detectedDirect disk access (anti-forensics, MFT access)
10ProcessAccessCredential dumping (LSASS access), process injection
11File createdRansomware file extension changes, data staging
12/13/14Registry eventsPersistence (Run keys, services), configuration changes
15FileCreateStreamHashAlternate Data Streams (ADS), malware hiding
17/18Pipe eventsNamed pipe usage (IPC, Cobalt Strike)
19/20/21WMI eventsWMI persistence, WMI-based lateral movement
22DNS queryC2 via domain, DGA detection
23File deletedFile deletion (ransomware note, cleanup)
25Process tamperingProcess hollowing, herpaderping
26File deleted (logged)Detailed file deletion logging
27FileBlockExecutableBlock executable file creation (if configured)
28FileBlockShreddingBlock file shredding (if configured)
29FileExecutableDetectedExecutable file detected (for detection)

Sysmon Configuration (Best Practice, SwiftOnSecurity):

<!-- sysmonconfig-export.xml by SwiftOnSecurity -->
<Sysmon schemaversion="4.90">
  <HashAlgorithms>sha256,IMPHASH</HashAlgorithms>
  <EventFiltering>
    <!-- Process creation -->
    <RuleGroup name="Process Creation" groupRelation="or">
      <ProcessCreate onmatch="include">
        <Image condition="begin with">C:\Windows\</Image>
        <Image condition="begin with">C:\Users\</Image>
        <CommandLine condition="contains">powershell</CommandLine>
        <CommandLine condition="contains">certutil</CommandLine>
        <CommandLine condition="contains">bitsadmin</CommandLine>
      </ProcessCreate>
    </RuleGroup>
    
    <!-- Network connections -->
    <RuleGroup name="Network Connections" groupRelation="or">
      <NetworkConnect onmatch="include">
        <DestinationPort condition="is">443</DestinationPort>
        <DestinationPort condition="is">80</DestinationPort>
        <DestinationPort condition="is">53</DestinationPort>
      </NetworkConnect>
    </RuleGroup>
    
    <!-- Registry events -->
    <RuleGroup name="Registry Events" groupRelation="or">
      <RegistryEvent onmatch="include">
        <TargetObject condition="contains">Run</TargetObject>
        <TargetObject condition="contains">RunOnce</TargetObject>
        <TargetObject condition="contains">Shell</TargetObject>
      </RegistryEvent>
    </RuleGroup>
    
    <!-- File creation -->
    <RuleGroup name="File Creation" groupRelation="or">
      <FileCreate onmatch="include">
        <TargetFilename condition="end with">.exe</TargetFilename>
        <TargetFilename condition="end with">.dll</TargetFilename>
        <TargetFilename condition="end with">.ps1</TargetFilename>
      </FileCreate>
    </RuleGroup>
    
    <!-- Process access (credential dumping) -->
    <RuleGroup name="Process Access" groupRelation="or">
      <ProcessAccess onmatch="include">
        <TargetImage condition="is">C:\Windows\system32\lsass.exe</TargetImage>
        <GrantedAccess condition="contains">0x1010</GrantedAccess>
      </ProcessAccess>
    </RuleGroup>
  </EventFiltering>
</Sysmon>

EDR + SIEM + NDR Correlation

The power of modern security monitoring comes from correlating endpoint, network, and cloud data.

Correlation Example: Detecting Ransomware

TimeSourceEventCorrelation Value
T+0EDRSysmon Event ID 1: powershell.exe executes vssadmin delete shadowsInitial indicator
T+30sEDRSysmon Event ID 11: Mass file creation with .locked extensionConfirms encryption
T+1mEDRNetworkConnect: Outbound connection to Tor exit nodePossible ransom payment channel
T+2mNetworkDNS query: payment-portal.onion.toC2 / payment infrastructure
T+5mCloudCloudTrail: IAM key created from compromised EC2 instanceLateral movement to cloud
T+10mEndpointFileCreateStreamHash: Ransom note README.txt in 100 foldersImpact confirmation

Correlation Query (Splunk):

| tstats `summariesonly` count from datamodel=Endpoint.Processes where Processes.process_name=vssadmin.exe by _time, host, user
| join host [
    | tstats `summariesonly` count from datamodel=Endpoint.Filesystem where Filesystem.file_name=*.locked by _time, host
]
| join host [
    | tstats `summariesonly` count from datamodel=Network_Resolution.DNS where DNS.query=*.onion.to by _time, host
]
| stats count by host, user
| where count > 3
| eval alert_type="ransomware_correlation"

XDR: Extended Detection and Response

XDR extends EDR by integrating data from multiple sources into a unified detection and response platform.

XDR PlatformIntegrated Data SourcesUnique Strength
Microsoft Defender XDREndpoint (MDE) + Identity (MDI) + Email (MDO) + Cloud Apps (MDCA) + SIEM (Sentinel)Native Microsoft ecosystem, unified incident scoring
CrowdStrike Falcon XDREndpoint + Identity + Cloud + Threat IntelligenceBest threat intel, Falcon Fusion SOAR
Palo Alto Cortex XDREndpoint + Firewall + Cloud + Threat IntelligenceNetwork + endpoint correlation, XDR agent + firewall integration
Trend Micro Vision OneEndpoint + Email + Network + Cloud + ServerGood breadth, XDR across email and network
SentinelOne VigilanceEndpoint + Cloud + Identity + NetworkAutonomous response, Storyline correlation

XDR vs. SIEM + EDR:

AspectSIEM + Separate EDRXDR
IntegrationManual correlation, API integrationNative, built-in correlation
Data normalizationRequires custom parsingVendor-normalized
Detection speedDepends on SIEM ingestion latencyTypically faster (native pipeline)
ResponseSOAR playbooks or manualIntegrated response across all vectors
Vendor lock-inLower (mix and match)Higher (single vendor ecosystem)
overheadVariable (multiple licenses)Often bundled, can be efficient
FlexibilityHigher (custom rules, any data source)Lower (vendor-controlled schema and rules)

💡 Tip from Singahi: For Microsoft-heavy organizations (Azure, M365, Windows), Microsoft Defender XDR + Sentinel is the most efficient and integrated option. For heterogeneous environments (Linux, AWS, GCP, diverse endpoints), a best-of-breed approach (CrowdStrike + Splunk/Elastic) provides more flexibility. The key is correlation, if your EDR and SIEM don't talk to each other, you're doing detection in silos.


Alert Management

Figure · Tiers

Maturity levels for monitoring activities

  1. InformationalAny
  2. LowLow
  3. MediumMedium
  4. HighHigh
  5. CriticalCritical
Where most organisations sit, and what the next level asks for. Full characteristics per level are in the table below.

A SIEM without alert management is a noise generator. Alert management is the discipline of turning raw alerts into actionable incidents.

Alert Triage Process

┌─────────────┐    ┌─────────────┐    ┌─────────────┐    ┌─────────────┐
│   ALERT     │ →  │  ENRICHMENT │ →  │  PRIORITIZE │ →  │  DECISION   │
│  (Raw SIEM  │    │  (Add context│    │  (Score risk,│    │  (Investigate│
│   detection) │    │  : user, asset│    │  check SLA) │    │  , escalate,│
│              │    │  , threat intel│    │             │    │  or dismiss)│
└─────────────┘    └─────────────┘    └─────────────┘    └─────────────┘

Alert Enrichment (Automated):

Enrichment TypeData SourceValue
User identityAD, Azure AD, HR systemIs this a VIP? Is this a terminated employee?
Asset criticalityCMDB, asset inventoryIs this a domain controller or a test server?
Threat intelligenceMISP, VirusTotal, CrowdStrikeIs this IP/domain/hash known malicious?
GeolocationMaxMind, IPinfoWhere is this IP located? Is it expected?
Vulnerability dataTenable, QualysIs this asset vulnerable to the detected technique?
Recent incidentsCase management systemHas this user/asset been in an incident before?
Peer group baselineUEBAIs this activity unusual for this user's role?

Alert Prioritization Framework

Not all alerts are equal. Use a risk-based prioritization matrix:

SeverityAsset CriticalityThreat CertaintyResponse SLAExample
CriticalCritical (DC, DB, CEO)Confirmed (IOC match, known malware)15 minutesRansomware on file server, DCSync detected
HighHigh (Production app, sensitive data)Likely (behavioral match + anomaly)1 hourImpossible travel + off-hours + bulk data access
MediumMedium (Standard workstation)Possible (single anomaly)4 hoursFailed login from unusual country
LowLow (Test environment, public data)Unlikely (policy violation only)24 hoursUSB device inserted on non-sensitive workstation
InformationalAnyContext onlyNo SLABaseline deviation within normal range

Alert Fatigue Reduction

Alert fatigue is the #1 reason SOC analysts miss real threats. The average SOC receives 10,000+ alerts per day. 90% are false positives.

Alert Fatigue Reduction Strategies:

StrategyImplementationExpected Reduction
Rule tuningExclude known benign patterns, refine thresholds30–50%
Alert suppressionSuppress alerts during maintenance windows, for known service accounts10–20%
Alert correlationGroup related alerts into single incident (e.g., 100 failed logins = 1 brute force incident)40–60%
Threshold elevationIncrease thresholds for low-value alerts after 30 days of baseline20–30%
Risk-based scoringOnly alert on high-risk scores; queue medium/low for batch review50–70%
ML-based filteringUse ML to classify alert likelihood as true positive20–40%
Automated enrichmentAuto-dismiss alerts where enrichment shows benign context10–20%
DeduplicationRemove duplicate alerts from multiple sensors (e.g., IDS + EDR both detect same malware)10–15%

Alert Correlation Example:

Before correlation:
- 10:00: Failed login from 192.168.1.1 (Alert #1)
- 10:01: Failed login from 192.168.1.1 (Alert #2)
- 10:02: Failed login from 192.168.1.1 (Alert #3)
... (50 alerts)
- 10:15: Successful login from 192.168.1.1 (Alert #51)
- 10:16: Privileged command executed (Alert #52)

After correlation:
- Incident #1: Brute force attack followed by successful compromise (1 incident, 52 correlated alerts)

SOAR, Security Orchestration, Automation, and Response

SOAR automates repetitive SOC tasks.

SOAR PlatformKey StrengthBest Forlicensing
Ansible + custom scriptsFully customizableTechnical teams, specific workflowsFree (labor overhead)

Common SOAR Playbooks:

PlaybookTriggerActionsTime Saved
Phishing responseEmail reported as phishingExtract indicators, check reputation, quarantine similar emails, block sender, create case30 min → 2 min
Malware containmentEDR detects malwareIsolate endpoint, block hash at firewall, search for hash across environment, create case45 min → 3 min
Account compromiseImpossible travel alertDisable account (with approval), force password reset, revoke sessions, notify manager, create case30 min → 2 min
Alert enrichmentAny high-priority alertEnrich IP with threat intel, check asset criticality, check user history, attach to case10 min → 30 sec
False positive feedbackAnalyst marks false positiveUpdate suppression rule, notify detection engineer, track FP rate15 min → 0 min

💡 Tip from Singahi: Alert fatigue is the silent killer of security monitoring programs. We've seen SOC teams with 50,000 daily alerts where analysts spend 100% of their time on false positives and 0% on real threats. If your SOC has more than 100 actionable alerts per analyst per day, you have an alert fatigue problem. The fix is not hiring more analysts, it's tuning your rules. We typically reduce alert volume by 60–80% in the first 30 days of engagement through systematic tuning.


The SOC Operating Model

A Security Operations Center (SOC) is the organizational engine that executes monitoring. Without a defined operating model, even the best tools fail.

SOC Tiers

TierRoleResponsibilityTypical ExperienceEscalation Trigger
L1, TriageSOC AnalystAlert triage, initial enrichment, false positive dismissal, playbook execution0–2 yearsConfirmed threat, needs investigation beyond playbook
L2, InvestigationSOC Analyst / Senior AnalystDeep investigation, threat hunting, malware analysis, IOC extraction2–5 yearsAPT, insider threat, major incident, needs incident response
L3, ExpertSenior Analyst / Threat HunterAdvanced forensics, detection engineering, threat intel, red team coordination5+ yearsNation-state, ransomware, major breach, needs executive briefing
L4, Incident ResponseIR Lead / CISOCrisis management, legal coordination, communications, post-incident review10+ yearsConfirmed breach, data exfiltration, regulatory notification required

SOC Shift Schedules

ModelCoverageStaff RequiredBest ForFatigue Risk
8×5 (Business hours)Mon–Fri, 9 AM–5 PM2–3 analystsSmall orgs, low risk, compliance-only SOCLow
12×5 (Extended)Mon–Fri, 7 AM–7 PM4–5 analystsMid-size, moderate riskLow
24×5 (Weekday always)Mon–Fri, 24 hours6–8 analystsMid-size, after-hours riskMedium
24×7 (Follow-the-sun)24/7 with regional handoffs9–12+ analystsLarge orgs, high risk, global operationsLow (if distributed)
24×7 (In-house shifts)24/7 with night shifts12–15 analystsLarge orgs, single geographyHigh (burnout risk)
24×7 (Hybrid: in-house + MDR)Business hours in-house, nights/weekends outsourced4–6 in-house + MDR contractMost growing companiesLow
Fully Managed (MDR)24/7 by vendor0–2 internalSmall orgs, no SOC capabilityLow (but dependency risk)

Shift Handoff Protocol:

Every shift change must include:

  1. Active incident status (open incidents, severity, current owner)
  2. Alerts requiring follow-up (triage incomplete, awaiting enrichment)
  3. System health issues (SIEM lag, agent failures, tool outages)
  4. Threat intelligence updates (new IOCs, active campaigns)
  5. Scheduled maintenance or changes (planned downtime, rule deployments)

SOC Runbooks

A runbook is a step-by-step procedure for handling specific alert types.

Runbook Structure:

Template

Escalation Matrix

ScenarioL1 ActionL2 ActionL3 ActionL4 ActionTimeline
Single malware alertTriage, playbookVerify, IOC extractionDetection updateNotification1 hour
Account compromiseIsolate, disableInvestigation, scopeThreat huntingIf customer data involved2 hours
Lateral movementImmediate escalationContainment, scopingFull incident leadCrisis management30 minutes
RansomwareImmediate escalationContainment coordinationTechnical leadBusiness continuity, legal15 minutes
Data exfiltrationImmediate escalationInvestigation, quantificationForensic leadRegulatory, legal, PR1 hour
Insider threatEscalate immediatelyCovert investigationForensic leadHR, legal, law enforcement2 hours

SOC Metrics & Staffing Model

Staffing Formula:

Analysts needed = (Alerts per day × Avg time per alert) / (Hours per day × Analyst efficiency)

Example:
- 2,000 alerts/day
- 10 minutes average triage time
- 8 hours/day effective work (accounting for breaks, meetings)
- 70% efficiency (30% of time spent on non-alert work: training, tuning, projects)

Analysts needed = (2,000 × 10) / (480 × 0.70) = 20,000 / 336 = 59 analyst-hours/day = 7.4 → 8 analysts

For 24/7 coverage with 8 analysts: Use hybrid MDR model or accept 8×5 coverage with on-call escalation.

Metrics & KPIs

What gets measured gets managed. Security monitoring metrics must be precise, actionable, and benchmarked.

The Essential Metrics Dashboard

MetricDefinitionTargetHow to Calculate
MTTD (Mean Time to Detect)Time from attack start to alert generation< 24 hoursAverage of (Alert time – Attack start time) across incidents
MTTR (Mean Time to Respond)Time from alert generation to initial containment< 1 hour for CriticalAverage of (Containment time – Alert time) across incidents
MTTC (Mean Time to Contain)Time from attack start to full containment< 4 hours for CriticalAverage of (Containment time – Attack start time)
Alert volumeTotal alerts per dayDeclining trend (tuning)Count of alerts in SIEM
False positive rate% of alerts that are false positives< 5%(False positives / Total alerts) × 100
Detection coverage% of MITRE ATT&CK techniques covered by rules> 70% for Tactics of interest(Covered techniques / Total techniques in scope) × 100
Detection rule countTotal active detection rulesGrowing trendCount in detection-as-code repo
Rule efficacy% of rules that generated true positives in last 30 days> 30%(Rules with TP / Total rules) × 100
SIEM ingestion rateGB/day or EPSWithin budget, stableSIEM dashboard or log aggregator metrics
Log source coverage% of critical assets sending logs100%(Assets forwarding logs / Total critical assets) × 100
Log retention compliance% of log sources meeting retention policy100%Audit storage vs. policy requirements
SLA compliance% of alerts meeting response SLA> 95%(Alerts within SLA / Total alerts) × 100
Incident closure rate% of incidents closed within target time> 90%(Closed on time / Total closed) × 100
SOC analyst use% of analyst time on value-added work70–80%(Time on investigation + hunting + tuning) / Total time
Threat hunting hoursHours per week spent on proactive hunting> 20% of SOC timeTracked via timesheet or ticket tagging
Mean time to tuneDays from rule deployment to acceptable FP rate< 14 daysAverage of (Tuning complete date – Deployment date)
SOAR automation rate% of alerts handled without human intervention30–60%(Auto-resolved alerts / Total alerts) × 100
Escalation rate% of alerts escalated to higher tier< 10%(Escalated alerts / Total alerts) × 100
Mean time to patch detection gapDays from new threat emergence to detection rule< 7 daysAverage for new high-priority threats

Benchmarking Your SOC

MetricUnderperformingAverageHigh-PerformingElite
MTTD> 7 days24–72 hours1–12 hours< 1 hour
MTTR> 4 hours1–4 hours15–60 minutes< 15 minutes
False positive rate> 20%10–20%5–10%< 5%
Alert volume per analyst/day> 200100–20050–100< 50
Detection coverage (ATT&CK)< 30%30–50%50–70%> 70%
SOAR automation rate< 10%10–30%30–50%> 50%
Mean time to patch gap> 30 days14–30 days7–14 days< 7 days

SIEM Performance Metrics

MetricWhat It MeansTargetAction if Failing
Ingestion latencyTime from log generation to SIEM availability< 5 minutesCheck network, collector health, parsing complexity
Query response timeTime for SIEM search to return results< 30 seconds for 24hAdd indexing, optimize queries, scale infrastructure
Indexer lagTime between real-time and indexed data< 1 minuteScale indexers, check resource constraints
Forwarder health% of log forwarders reporting healthy100%Monitor forwarder status, auto-restart failed agents
Disk use% of storage used< 80%Archive old data, add storage, reduce retention
License use% of SIEM license used< 85% (headroom)Filter noise, optimize ingestion, renegotiate license
Search head concurrencyNumber of simultaneous searchesWithin capacityAdd search heads, optimize scheduled searches

Tool Comparison

Complete SIEM / Security Monitoring Matrix

ToolTypeDeploymentlicensingBest ForSIEMUEBASOAREDR IntegrationCloud NativeOpen Source
WazuhSIEM/XDRSelf-managedFreeBudget-conscious, compliance✅⚠️⚠️✅ (agent)⚠️✅
ArkimePCAP/NDRSelf-managedFreeFull packet capture, forensics❌❌❌⚠️⚠️✅
ZeekNSM/NDRSelf-managedFreeNetwork analysis, protocol detection❌❌❌⚠️⚠️✅
SuricataIDS/IPSSelf-managedFreeNetwork intrusion detection❌❌❌⚠️⚠️✅
FalcoContainer RuntimeSelf-managedFreeContainer security, K8s runtime❌❌⚠️⚠️✅ (K8s)✅

Open Source Stack for Budget-Conscious Organizations

If budget is constrained, this open-source stack provides 80% of enterprise capability at 10% of the overhead:

LayerToolRoleoverhead
Log collectionFluent Bit + FilebeatLightweight log shippersFree
Log aggregationKafka or syslog-ngBuffering and routingFree
SIEMElastic Security (self-managed)Core SIEM, detection, casesFree (open source)
Endpoint monitoringWazuh + SysmonEDR-like capability, FIM, rootkit detectionFree
Network monitoringSuricata + ZeekIDS + network analysisFree
Container monitoringFalco + K8s auditContainer runtime securityFree
Threat huntingJupyter + Python + ElasticCustom hunting notebooksFree
VulnerabilityOpenVAS + WazuhVulnerability scanningFree
Case managementTheHive + CortexCase management + IOC analysisFree
AutomationShuffle or custom scriptsSOAR-liteFree

Trade-offs:

  • No vendor support (community support only)
  • Requires internal expertise
  • Higher time-to-value
  • Less polished UX
  • Integration between tools requires manual work

Implementation Roadmap: 8 Weeks

Figure · Timeline

Escalation timeline

  1. Day 1Inventory all critical assets
  2. Day 2Map existing log sources
  3. Day 3Define monitoring policy
  4. Day 4Select SIEM/tooling
  5. Day 5Design architecture
  6. 6–7Deploy basic log collection
Milestones in delivery order. Owners and the evidence each produces are in the table below.

Week 1: Assessment & Planning

DayActivityDeliverableOwner
1Inventory all critical assets (servers, apps, DBs, network devices, cloud)Asset inventory with criticality ratingsIT Manager
2Map existing log sources (what's already logging, what's missing)Log source gap analysisSecurity Analyst
3Define monitoring policy (scope, retention, review frequency, privacy)Draft monitoring policyCISO
4Select SIEM/tooling (based on budget, org size, tech stack)Tool selection matrix with recommendationCISO + Procurement
5Design architecture (collectors, forwarders, network diagram)Architecture diagramSecurity Engineer
6–7Deploy basic log collection (syslog server, Windows Event Forwarding)Central log collector operationalSecurity Engineer

Week 2: Core Infrastructure

DayActivityDeliverable
8–9Deploy SIEM (install, license, basic configuration)SIEM accessible, ready for ingestion
10–11Configure critical log sources (firewall, AD, key servers)5+ log sources forwarding to SIEM
12Implement time synchronization (NTP across all sources)NTP configured, drift < 1 second
13Configure log parsing and normalization (ECS or CIM)3+ log sources normalized
14Test log flow end-to-end (source → collector → SIEM)Validation report

Week 3: Detection Rules

DayActivityDeliverable
15–16Deploy top 10 detection rules (failed logins, privilege escalation, malware)10 rules active, tested
17–18Configure basic alerting (email, Slack, PagerDuty)Alert channel operational
19Test detections with known bad events (simulate failed logins, test account)Detection test report
20Configure log retention and archiving (hot/warm/cold tiers)Retention policy implemented
21Document runbooks for top 5 alert types5 runbooks published

Week 4: Cloud & Advanced Sources

DayActivityDeliverable
22–23Enable cloud audit logs (CloudTrail, Azure Monitor, GCP audit)Cloud logs flowing to SIEM
24Deploy endpoint monitoring (EDR or Wazuh)80% of endpoints reporting
25Configure database audit logs (top 3 critical databases)DB audit enabled, forwarded
26Enable network flow monitoring (NetFlow, VPC Flow Logs)Flow data visible in SIEM
27–28Configure K8s audit logs and container runtime (if applicable)Container monitoring active

Week 5: Alert Management & Triage

DayActivityDeliverable
29–30Implement alert enrichment (threat intel, asset criticality)Enrichment pipeline active
31–32Deploy SOAR playbooks (phishing, malware, account compromise)3 playbooks automated
33Configure alert suppression and deduplication20% alert volume reduction
34Establish SOC shift schedule and handoff protocolSchedule published, training complete
35–36Conduct SOC training (tool usage, runbooks, escalation)All analysts trained

Week 6: UEBA & Behavioral Analytics

DayActivityDeliverable
37–38Deploy UEBA (built-in SIEM UEBA or standalone tool)UEBA baseline learning started
39–40Configure behavioral baselines (login times, data access, geolocation)Baselines established for 50% of users
41Implement risk scoring modelRisk scores visible in SIEM
42Test UEBA with simulated anomaly (impossible travel test)UEBA detection confirmed

Week 7: Detection Engineering & Tuning

DayActivityDeliverable
43–44Expand detection rules to 50 (cover top 20 ATT&CK techniques)50 rules active
45Map rules to MITRE ATT&CKATT&CK coverage matrix published
46–47Tune false positives (target < 10% FP rate)Tuning report, FP rate measured
48Implement detection-as-code repository (Git + CI/CD)Detection repo operational

Week 8: Metrics, Audit Readiness & Handoff

DayActivityDeliverable
49–50Build metrics dashboard (MTTD, MTTR, alert volume, coverage)Dashboard live
51Conduct internal audit (review logs, test alerts, verify retention)Internal audit report
52Document evidence for ISO 27001 audit (policy, procedures, records)Evidence package ready
53–54Final tuning and optimization (SIEM performance, query speed)Performance report
55–56Project handoff and operational transition (SOC takes ownership)Handoff document signed

Week 8 Milestone: ISO 27001 A.8.16 Audit Ready

Audit RequirementEvidence
Monitoring policyApproved, published, communicated
Log source coverage100% of critical assets confirmed
Centralized collectionSIEM dashboard showing live ingestion
Detection rules50+ rules mapped to risk register
Alert review records2 weeks of documented alert triage
Anomaly responseRunbooks with escalation procedures
Retention compliance90-day hot, 1-year warm, 3-year cold
Clock synchronizationNTP configured, drift verified
MetricsDashboard showing MTTD, MTTR, coverage

Common Audit Failures & Fixes

After guiding 50+ companies through ISO 27001 audits, we've seen the same A.8.16 non-conformities repeatedly. Here's how to avoid them.

Failure 1: No Centralized Log Collection

Auditor finding: "Logs are stored only on local systems. There is no evidence of centralized collection or review."

Why it happens: Organizations enable logging on each server but never configure forwarding. They think "we have logs" means "we are compliant."

The fix:

  1. Deploy a syslog collector (rsyslog, syslog-ng, or Splunk UF) within 2 weeks
  2. Forward all critical Windows events via WEF or Winlogbeat
  3. Forward all Linux auth logs via rsyslog
  4. Verify collection daily: create a dashboard showing log source health
  5. Document the architecture in the monitoring policy

Evidence for audit: SIEM dashboard showing active log sources, network diagram of log flow, policy section on collection.

Failure 2: Logs Are Not Reviewed

Auditor finding: "Logs are collected but there is no evidence of regular review or analysis for security events."

Why it happens: The SIEM is a "set and forget" tool. Alerts fire into email but no one has time to review them.

The fix:

  1. Define review frequency: real-time alerts for critical events, daily triage for high/medium, weekly trend analysis
  2. Assign specific people to review (named in policy, not "the IT team")
  3. Document every review: date, reviewer, scope, findings, actions taken
  4. Start with just 30 minutes per day. Review yesterday's authentication failures and privilege changes.
  5. Use SOAR to auto-dismiss obvious false positives, reducing review burden

Evidence for audit: Log review logs, daily triage tickets, weekly summary reports, incident records created from log reviews.

Failure 3: No Alerting on Critical Events

Auditor finding: "The organization collects logs but has no automated alerting configured for anomalous behavior."

Why it happens: Organizations confuse log collection with monitoring. Monitoring requires detection and response.

The fix:

  1. Define 10 critical events that must generate alerts (see Section 8)
  2. Configure these alerts in your SIEM within 1 week
  3. Ensure alerts go to someone who is accountable (not a group mailbox that no one owns)
  4. Test each alert monthly: simulate the event, verify the alert fires
  5. Document alert configuration and response procedures

Evidence for audit: Screenshot of alert configuration, test records, incident tickets created from alerts.

Failure 4: Insufficient Log Retention

Auditor finding: "Logs are retained for 30 days, but the policy requires 90 days, and regulatory requirements mandate 1 year."

Why it happens: Default settings are accepted. Disk space is premium-tier. No one checks retention against policy.

The fix:

  1. Define retention periods by log type and regulation:
    • Authentication logs: 1 year
    • Firewall logs: 90 days hot, 1 year warm
    • Application logs: 90 days
    • Cloud audit logs: 1 year (AWS CloudTrail default)
  2. Implement hot/warm/cold tiering: recent data in fast storage, old data in cheap archive
  3. Automate archival: S3 Glacier, Azure Archive, or tape backup
  4. Verify retention quarterly: restore a random archive log and confirm readability
  5. Document retention policy and technical implementation

Evidence for audit: Retention policy, storage configuration, sample archived log, quarterly verification records.

Failure 5: Clock Synchronization Issues

Auditor finding: "System clocks are not synchronized. Log timestamps vary by more than 5 minutes across systems, making correlation impossible."

Why it happens: NTP is not configured on all devices, or devices are using different NTP servers with drift.

The fix:

  1. Deploy central NTP servers (pool.ntp.org or internal Stratum 2)
  2. Configure all Windows systems to sync with domain controllers (which sync to NTP)
  3. Configure all Linux systems with chronyd or ntpd pointing to central NTP
  4. Configure network devices, firewalls, and appliances to sync to NTP
  5. Verify daily: run a script that checks time drift across all systems
  6. Set alarm if drift > 1 second

Evidence for audit: NTP configuration files, time drift check results, SIEM query showing consistent timestamps.

Failure 6: No Monitoring of Cloud Infrastructure

Auditor finding: "The organization operates workloads in AWS but the monitoring policy and SIEM do not cover cloud events."

Why it happens: Cloud monitoring requires different tools and expertise. On-prem teams don't extend their scope to cloud.

The fix:

  1. Enable CloudTrail (AWS), Activity Log (Azure), or Cloud Audit Logs (GCP) in all regions
  2. Forward cloud audit logs to your SIEM using native connectors or Event Hub/Kinesis
  3. Configure alerts for cloud-specific events: IAM changes, public resource creation, root account use
  4. Include cloud in monitoring policy scope
  5. Assign cloud security monitoring responsibility to a named person

Evidence for audit: CloudTrail configuration screenshot, SIEM showing cloud events, policy section on cloud monitoring.

Failure 7: Privileged User Activity Not Monitored

Auditor finding: "There is no specific monitoring or alerting for administrator and privileged user activities."

Why it happens: Administrators are trusted, so their activity is not scrutinized. Or, admins object to being monitored.

The fix:

  1. Log all privileged actions: sudo usage, admin group membership changes, IAM policy changes
  2. Alert on privileged actions: create a dedicated alert channel for admin activity
  3. Review admin activity weekly: separate review from standard user activity
  4. Implement break-glass procedures: emergency admin access requires manager approval and is fully logged
  5. No admin should be able to clear or modify their own logs

Evidence for audit: Privileged user monitoring dashboard, weekly admin review records, break-glass procedure.

Failure 8: Monitoring Policy Does Not Exist

Auditor finding: "There is no documented policy or procedure for monitoring activities."

Why it happens: Organizations focus on tools and forget governance. "We have Splunk" is not a policy.

The fix:

  1. Write a monitoring policy (see Section 2 for minimum requirements)
  2. Get it approved by management (CISO or CEO depending on org size)
  3. Publish it on the intranet or document management system
  4. Communicate it to all relevant personnel (IT, security, operations)
  5. Review annually and after significant changes

Evidence for audit: Signed policy document, communication records, acknowledgment tracking, review records.

Failure 9: No Response Procedure for Anomalies

Auditor finding: "Anomalies are detected but there is no defined procedure for evaluating or escalating potential incidents."

Why it happens: Alerts fire but no one knows what to do. The link between monitoring and incident response is broken.

The fix:

  1. Define anomaly severity levels (Critical, High, Medium, Low)
  2. Map each severity to a response action:
    • Critical: Immediate SOC investigation, potential incident escalation
    • High: SOC investigation within 1 hour
    • Medium: Daily batch review
    • Low: Weekly trend review
  3. Document escalation to incident response (A.5.24)
  4. Create runbooks for top 5 anomaly types (see Section 17)
  5. Test the escalation path quarterly: simulate an anomaly, verify IR is engaged

Evidence for audit: Anomaly response procedure, runbooks, escalation test records, incident tickets created from anomalies.

Failure 10: Third-Party Access Not Monitored

Auditor finding: "Vendor and contractor access to systems is not subject to the same monitoring as employee access."

Why it happens: Third-party accounts are created but not included in monitoring scope. Vendor VPNs are separate from corporate VPN.

The fix:

  1. Include all third-party accounts in identity monitoring (same as employees)
  2. Log and alert on all vendor access: VPN, RDP, cloud console, application access
  3. Review vendor access weekly or monthly
  4. Implement time-bound access: vendor accounts expire automatically
  5. Require vendors to provide their own audit logs for critical systems

Evidence for audit: Vendor access logs, vendor monitoring dashboard, third-party access review records.


Illustrative Scenarios: Breaches That Monitoring Could Have Prevented

Illustrative scenario, a composite example for guidance, not a specific Singahi engagement or a verified outcome.

What happened:

  • Attackers compromised an HVAC vendor's credentials (Fazio Mechanical)
  • Used vendor access to enter Target's network
  • Moved laterally to the POS system
  • Installed malware on POS terminals
  • Exfiltrated 40 million credit card numbers and 70 million customer records

Monitoring failures:

  • Target had a SIEM (FireEye) but did not monitor vendor access adequately
  • The SIEM generated alerts for the malware installation, but alerts were not reviewed or acted upon
  • Lateral movement (HVAC network → POS network) was not segmented or monitored
  • Data exfiltration over weeks went undetected

What proper monitoring would have caught:

  1. Vendor access alert: "Vendor account logged in at 2 AM, accessed file server, then POS network"
  2. Malware detection: FireEye alert (which did fire) escalated to immediate response
  3. Lateral movement: "Connection from HVAC VLAN to POS VLAN detected"
  4. Data exfiltration: "Unusual outbound FTP connection from POS server"

Lesson: A SIEM without alert review and response is just premium-tier storage. The FireEye alert fired. No one acted.

Illustrative Scenario 2: Equifax (2017), The 147 Million Record Breach

What happened:

  • Apache Struts vulnerability (CVE-2017-5638) was exploited
  • Attackers gained access to Equifax's dispute portal
  • Moved laterally through the network
  • Accessed databases containing personal data
  • Exfiltrated data over 76 days

Monitoring failures:

  • Vulnerability scanning (A.8.8) failed to identify the Apache Struts vulnerability promptly
  • Network monitoring did not detect the initial exploitation or lateral movement
  • Database monitoring did not detect unauthorized database queries
  • Data exfiltration went unnoticed for 76 days
  • SSL certificate monitoring failed: Equifax discovered the breach only when they renewed an SSL certificate and noticed anomalous traffic

What proper monitoring would have caught:

  1. Web application monitoring: "Apache Struts exploitation attempt detected by WAF"
  2. Network monitoring: "Unusual connection from web server to database server"
  3. Database monitoring: "Database queries from web server IP returning 147M records"
  4. Data exfiltration: "76-day sustained outbound transfer of 30GB+ to unknown IP"
  5. Certificate monitoring: This should not have been the detection mechanism

Lesson: Monitoring must be layered. Relying on a single point of failure (SSL certificate renewal) for breach detection is negligence. Network + application + database monitoring would have caught this at multiple stages.

Illustrative Scenario 3: SolarWinds / Sunburst (2020), The Supply Chain Breach

What happened:

  • Attackers compromised SolarWinds' build system
  • Inserted malicious code into Orion software updates
  • 18,000 organizations downloaded the backdoored update
  • Attackers used the backdoor to access select high-value targets (including government agencies)
  • Activity went undetected for months

Monitoring failures:

  • Organizations that did not monitor outbound DNS from their Orion servers missed the C2 beaconing
  • UEBA was missing: The Orion server making outbound HTTPS to unknown domains was anomalous
  • Supply chain monitoring: No one monitored the integrity of the software update itself
  • Cloud monitoring: Attackers accessed Azure AD and Office 365, which was not monitored for anomalous application access

What proper monitoring would have caught:

  1. DNS monitoring: "Orion server querying avsvmcloud.com [DGA-like domain] every 60 seconds"
  2. UEBA: "SolarWinds server making HTTPS to unknown domain, never seen before"
  3. Cloud monitoring: "Application 'SolarWinds Orion' accessing Azure AD with unusual permissions"
  4. Network monitoring: "Beaconing pattern detected: 60-second intervals to same external IP"

Lesson: Supply chain attacks require monitoring the vendor's behavior within your environment. The software was trusted, but its behavior was anomalous. DNS and behavioral monitoring catch supply chain compromises when signature-based tools cannot.

Illustrative Scenario 4: Colonial Pipeline (2021), Ransomware and Operational Shutdown

What happened:

  • Attackers compromised Colonial Pipeline via a leaked VPN password
  • Used the VPN to access the business network
  • Deployed DarkSide ransomware
  • Colonial shut down the pipeline (operational network) as a precaution
  • Caused fuel shortages across the US East Coast

Monitoring failures:

  • VPN monitoring: No alert for VPN login from unusual location or with old password
  • Credential monitoring: The VPN password was discovered in a breach database (no credential monitoring)
  • OT/ICS monitoring: Separation between IT and OT was assumed but not monitored
  • Ransomware indicators: No detection of ransomware precursor activities (AD reconnaissance, lateral movement, backup deletion)

What proper monitoring would have caught:

  1. VPN monitoring: "VPN login from unexpected geography using old password"
  2. Credential monitoring: "Password for VPN account found in Have I Been Pwned breach database"
  3. Network monitoring: "Connection from IT network to OT network detected"
  4. Endpoint monitoring: "DarkSide ransomware behavior: vssadmin delete shadows, mass file encryption"

Lesson: OT/ICS monitoring is critical. The pipeline was shut down not because ransomware hit the OT network, but because Colonial could not confirm that the OT network was safe. Monitoring the IT/OT boundary would have provided that confidence.

Illustrative Scenario 5: Uber (2016), The Cover-Up

What happened:

  • Attackers accessed Uber's GitHub repository and found AWS access keys
  • Used keys to access Uber's AWS S3 buckets
  • Downloaded 57 million user and driver records
  • Disclosure was delayed for over a year

Monitoring failures:

  • Cloud monitoring: No alert for AWS access key usage from unknown IP or for bulk S3 downloads
  • Code repository monitoring: No monitoring of GitHub for exposed credentials (secrets scanning)
  • Data access monitoring: No alert for 57 million record download from S3
  • Third-party monitoring: GitHub is a third-party SaaS; Uber did not monitor it for security

What proper monitoring would have caught:

  1. Cloud monitoring: "AWS access key used from unknown IP address"
  2. Cloud monitoring: "S3 bulk download: 57M objects accessed from single session"
  3. Secrets monitoring: "AWS access key committed to GitHub repository"
  4. Data monitoring: "Unusual data access pattern: entire user database downloaded"

Lesson: Cloud-native monitoring is essential. AWS CloudTrail would have recorded every API call. The data was there, Uber was not watching it. Additionally, secrets scanning in CI/CD (A.8.30) is a prerequisite to monitoring, don't commit credentials to repositories.


Multi-Framework Mapping

A.8.16 is not an island. Mapping it to other frameworks ensures compliance efficiency and audit readiness.

Annex A 8.16 Across Major Frameworks

FrameworkControl ReferenceEquivalent RequirementKey Difference
ISO 27001:2022A.8.16Monitor networks, systems, and applications for anomalous behaviorBroadest scope; includes "appropriate actions"
ISO 27001:2013A.12.4.1, A.12.4.2, A.12.4.3Event logging, protection of logs, admin/operator logs2022 merges and refocuses on active monitoring
SOC 2 (TSC 2017)CC7.2System monitoringRequires monitoring of system components; less prescriptive on tools
PCI DSS v4.0Req 10.1–10.7Logging and monitoringFocuses on CHD environment; requires specific log review (daily for critical systems)
NIST SP 800-53 Rev 5AU-6, AU-12, IR-4Audit review, audit generation, incident handlingTechnical depth; requires automated tools for high-impact systems
NIST CSF 2.0DE.AE-1, DE.AE-2, DE.AE-3Anomaly detection, event detection, log analysisOutcome-focused; doesn't specify tools
DORA (EU)Art. 10(2), Art. 11ICT risk management, incident reportingRequires continuous monitoring for financial entities; mandatory reporting
COBIT 2019DSS05.05, APO12.06Monitor and review security controlsGovernance-focused; links to risk management
ITIL 4Practice: Information SecuritySecurity monitoring as part of service managementService-focused, less technical
CIS Controls v8Controls 8, 13, 16Audit logging, network monitoring, application loggingPrescriptive technical controls; specific implementation guidance
GDPR (EU)Art. 32, Art. 33Security of processing, breach notificationMonitoring must support breach detection and notification within 72 hours
HIPAA (US)164.312(b)Audit controlsMust record and examine access to ePHI; requires regular review

Mapping: ISO 27001 A.8.16 → SOC 2 CC7.2

ISO 27001 A.8.16 ComponentSOC 2 CC7.2 Trust Service CriteriaEvidence Needed
Monitoring strategyCC7.2.1: System components monitoredMonitoring policy, scope documentation
Log collection and centralizationCC7.2.2: Event data capturedSIEM configuration, log source inventory
Anomaly detectionCC7.2.3: Anomalies detectedDetection rules, alert configuration
Log reviewCC7.2.4: Events analyzedReview logs, documented analysis
Response to anomaliesCC7.2.5: Incidents identifiedIncident tickets, escalation records
RetentionCC7.2.6: Audit logs retainedRetention policy, storage configuration
Clock synchronizationCC7.2.7: Consistent time sourcesNTP configuration, time drift checks

Mapping: ISO 27001 A.8.16 → PCI DSS v4.0

PCI DSS v4.0 RequirementA.8.16 ImplementationAdditional PCI Requirements
Req 10.1: Audit trailsEnable logging on all CHD systemsMust cover all user actions, including privileged users
Req 10.2: Audit trail contentConfigure detailed OS, application, database logsMust include user ID, event type, date/time, success/failure, event origin
Req 10.3: Clock syncNTP on all CHD systemsSynchronization must be accurate within 1 minute
Req 10.4: Log accessibilityCentralized collection with tamper protectionLogs must be immediately available and unmodifiable
Req 10.5: Log reviewDaily review of critical logs, weekly for othersDocumented review with evidence
Req 10.6: Time synchronizationSame as A.8.17Time must be consistent across all systems
Req 10.7: Retention1 year hot, 3 months immediately availableSpecific retention requirements for CHD

Mapping: ISO 27001 A.8.16 → NIST SP 800-53 Rev 5

NIST ControlTitleA.8.16 MappingImplementation Guidance
AU-6Audit Record ReviewAnomaly detection and log reviewAutomated review tools + manual review
AU-12Audit Record GenerationLog source configurationAll security-relevant events must be logged
AU-13Monitoring for Information DisclosureData exfiltration detectionDLP integration, network monitoring
IR-4Incident HandlingResponse to anomaliesSOC runbooks, escalation procedures
IR-5Incident MonitoringSOC operationsContinuous monitoring of incident status
SI-4Information System MonitoringNetwork and system monitoringIDS, SIEM, network behavior analysis
SI-4(1)System Monitoring, Intrusion DetectionIDS/IPS deploymentAutomated real-time intrusion detection
SI-4(2)System Monitoring, Automated AlertsAlert configurationAutomated alerts for critical events
SI-4(4)System Monitoring, Inbound/Outbound TrafficNetwork monitoringMonitor both inbound and outbound traffic
SI-4(5)System Monitoring, System-Generated AlertsSystem-level alertingOS-level alerts, not just network
SI-4(7)System Monitoring, Automated ResponseSOAR / automated responseAutomated containment actions

Mapping: ISO 27001 A.8.16 → DORA (EU Digital Operational Resilience Act)

DORA ArticleRequirementA.8.16 Implementation
Art. 10(2)Continuous monitoring of ICT systems24/7 monitoring for critical systems
Art. 10(3)Anomaly detectionUEBA, behavioral analytics, ML-based detection
Art. 11(1)Incident detection and reportingSOC with defined detection and escalation procedures
Art. 11(2)Incident classificationSeverity levels mapped to response SLAs
Art. 11(3)Major incident reportingAutomated reporting to regulators within 4 hours
Art. 12(1)Digital operational resilience testingRed team exercises, alert testing, SOC drills

FAQ

Q1: Do I need a SIEM to pass an ISO 27001 audit for A.8.16?

No, but you need centralized monitoring and anomaly detection. A SIEM is the most common and effective way to achieve this. For small organizations, a centralized syslog server with daily manual review and basic alerting may be sufficient. However, as you scale, a SIEM becomes essential. The auditor will check: Are you monitoring? Are you detecting anomalies? Are you responding? A SIEM makes this easier but is not strictly mandatory.

Q2: What's the minimum budget for A.8.16 compliance?

For a small organization (under 50 employees), the minimum budget is:

For mid-market (50–500 employees):

For enterprise (500+ employees):

Q3: How do I monitor cloud-native services (serverless, managed databases)?

Cloud-native services are monitored via the cloud provider's APIs:

  • Serverless (Lambda, Azure Functions, Cloud Functions): Enable function-level logging (CloudWatch Logs, Azure Monitor, Cloud Logging). Log every invocation, error, and cold start. Forward to SIEM.
  • Managed databases (RDS, Azure SQL, Cloud Spanner): Enable native audit logs (RDS audit, Azure SQL Auditing, Cloud Audit Logs). These are API-accessible and forwardable to SIEM.
  • Managed Kubernetes (EKS, AKS, GKE): Enable control plane logging (API server, audit, authenticator). Forward via Fluent Bit or cloud-native export.
  • API Gateways: Enable access logging and execution logging. Forward to SIEM or cloud monitoring.

The key is to not treat cloud services as "black boxes." Every cloud service has logs. You just need to enable them and forward them.

Q4: What retention period is required for logs?

ISO 27001 does not specify a fixed retention period. It depends on:

  • Your monitoring policy (which you define)
  • Regulatory requirements (PCI DSS: 1 year, GDPR: as long as necessary for security)
  • Legal requirements (litigation hold may require longer retention)
  • Industry standards (financial services often require 3–7 years)

Singahi recommendation:

  • Authentication logs: 1 year
  • Security event logs: 1 year
  • Application logs: 90 days
  • Network flow logs: 90 days
  • Full packet capture: 7–30 days (unless incident-related)
  • Cloud audit logs: 1 year (or as required by cloud provider)
  • Archive logs: 3–7 years in cold storage

Q5: How do I handle privacy concerns with user monitoring?

Monitoring user activity must balance security with privacy:

  • Document monitoring in your acceptable use policy and privacy policy
  • In EU, consult works council if applicable
  • Monitor access patterns, not content where possible (e.g., "accessed 500 files" not "read confidential merger memo")
  • Avoid keystroke logging or screen recording without explicit consent
  • Implement data minimization: only collect what is necessary for security
  • Set retention limits: delete user activity baselines when the user departs
  • Allow access to one's own logs for transparency (where legally required)

Q6: What's the difference between A.8.15 (Logging) and A.8.16 (Monitoring)?

  • A.8.15, Logging: The generation and protection of logs. Ensures logs exist, are accurate, and are protected from tampering. Technical control.
  • A.8.16, Monitoring: The analysis of logs and systems for anomalous behavior. Requires detection, review, and response. Operational control.

You need both. A.8.15 without A.8.16 means you have logs but no one looks at them. A.8.16 without A.8.15 means you want to monitor but have no logs to analyze.

Q7: How do I detect zero-day attacks?

Zero-day attacks cannot be detected by signatures (no one knows the signature yet). Detection relies on:

  • Behavioral monitoring: Anomalous process execution, unusual network connections, unexpected file changes (EDR, UEBA)
  • Network anomaly detection: Beaconing, unusual data flows, DGA domains (NDR, network monitoring)
  • Threat hunting: Proactive searches for indicators of unknown threats (e.g., "show me all processes that made outbound connections and were not seen in the last 30 days")
  • ML-based anomaly detection: UEBA and network behavior analytics that detect deviations from baseline
  • Honeypots / deception: Fake assets that attract attackers; any interaction is anomalous

Q8: Should I build an in-house SOC or use an MDR?

FactorIn-House SOCMDR (Managed Detection and Response)
overheadHigh (salaries, tools, training)Predictable monthly fee
ControlFull control over tools and processesVendor-controlled, limited customization
ExpertiseBuild over time, harder to hireImmediate access to experienced analysts
ScalabilityRequires hiring and trainingScales with contract
24/7 coveragepremium-tier to staff nights/weekendsIncluded in MDR service
CustomizationFully customizableLimited to vendor's capabilities
ContextDeep knowledge of your environmentLess context, requires knowledge transfer
Best forLarge orgs, regulated industries, unique environmentsSmall-mid orgs, no SOC expertise, budget-constrained

Singahi recommendation:

  • Under 250 employees: MDR (fully managed)
  • 250–1,000 employees: Hybrid (in-house business hours + MDR for nights/weekends)
  • 1,000+ employees: In-house SOC with MDR for overflow or specific functions (e.g., threat hunting)

Q9: How do I measure the effectiveness of my monitoring program?

Use the metrics from Section 18. The top 5 metrics are:

  1. MTTD (Mean Time to Detect): Are you finding threats faster? Target < 24 hours.
  2. MTTR (Mean Time to Respond): Are you responding faster? Target < 1 hour for critical.
  3. False positive rate: Are you wasting analyst time? Target < 5%.
  4. Detection coverage: Are you covered against the threats you care about? Target > 70% of ATT&CK techniques in scope.
  5. Alert volume per analyst: Is your SOC sustainable? Target < 100 actionable alerts per analyst per day.

Q10: What should I do if my SIEM is too premium-tier?

Options:

  1. Reduce ingestion: Filter noise, sample high-volume logs, drop debug logs
  2. Optimize storage: Use hot/warm/cold tiers; move old data to cheap archive
  3. Switch to open source: Elastic (self-managed) or Wazuh for core SIEM capability
  4. Use cloud-native SIEM: Sentinel or Chronicle can be efficient for cloud-first orgs
  5. Hybrid approach: Use cloud SIEM for real-time alerting and data lake (S3, BigQuery) for long-term storage and batch analysis
  6. Renegotiate: Many vendors offer significant discounts if you commit to volume tiers or multi-year contracts

Q11: How do I monitor third-party SaaS applications (Slack, Salesforce, GitHub)?

Most SaaS applications provide audit logs via API:

  • Slack: Audit logs (Enterprise Grid) via API
  • Salesforce: Event Monitoring, Setup Audit Trail
  • GitHub: Audit log, Security tab, Dependabot alerts
  • Google Workspace: Admin console audit, Google Cloud Identity logs
  • Office 365: Unified Audit Log (via Compliance Center or Sentinel connector)
  • Jira/Confluence: Audit log (Data Center/Cloud)
  • Zoom: Account activity, operation log

Collect these via API connectors (Splunk Add-on, Sentinel Data Connector, Elastic Agent) or custom scripts. Include SaaS in your monitoring policy scope.

Q12: Can I use my existing observability tool (Datadog, New Relic) for security monitoring?

Yes, with caveats:

  • Pros: Unified platform, existing infrastructure, overhead efficiency, cloud-native
  • Cons: Security features are newer and less mature than dedicated SIEMs; limited SOAR, limited detection rule libraries, less strong case management

Best approach:

  • Use observability for application security (API abuse, injection, performance-based anomalies)
  • Use dedicated SIEM for enterprise security monitoring (correlation, threat hunting, compliance)
  • Integrate both: forward observability security findings to SIEM for unified incident management

Q13: What is the most common mistake in A.8.16 implementation?

The most common mistake is deploying tools without process. Organizations buy Splunk, deploy agents, and consider the job done. They don't have:

  • A monitoring policy
  • Defined review procedures
  • Alert tuning and false positive management
  • Escalation to incident response
  • Metrics and KPIs

The tool is not the control. The process is the control.

Q14: How often should detection rules be updated?

Rule TypeUpdate FrequencyTrigger
IOC-based rules (hashes, IPs, domains)Daily / WeeklyNew threat intelligence feeds
Behavioral rules (UEBA, baselines)MonthlyBaseline drift, seasonal changes
TTP-based rules (ATT&CK techniques)QuarterlyNew ATT&CK versions, new attack research
Compliance rulesAnnuallyPolicy changes, audit findings
Red team feedbackAfter each exerciseNew red team TTPs, evasion techniques

Q15: Do I need to monitor OT/ICS environments for ISO 27001?

If your scope includes OT/ICS (manufacturing, energy, critical infrastructure), then yes. However, OT monitoring requires specialized tools and approaches:

  • Passive monitoring (no active scanning that could disrupt operations)
  • Protocol-aware (Modbus, DNP3, IEC 61850)
  • Air-gap considerations (data diode for log export)
  • Tools: Nozomi Networks, Claroty, Dragos, Fortinet OT Security

If your scope is purely IT, OT may be excluded with explicit scope definition and risk acceptance.

Indian Regulatory Context and Illustrative Scenario for A.8.16

Indian regulators expect security monitoring to be continuous, actionable and aligned with sectoral reporting obligations. RBI's Cyber Security Framework requires banks to deploy SIEM, correlate logs centrally, and report serious incidents within 6 hours. SEBI's cybersecurity circulars for market infrastructure institutions mandate 24x7 security operations, log retention for at least one year and regular vulnerability scanning. CERT-In directions require reporting of specified cyber incidents within 6 hours and maintenance of logs for 180 days. The DPDP Act 2023 makes monitoring evidence critical for demonstrating reasonable security safeguards (Section 8(5)) and for detecting personal data breaches that must be intimated under Section 8(6).

Illustrative Scenario, Indian E-commerce Fraud Detection (2024): A Gurugram-based e-commerce platform noticed anomalous order refunds totaling over a two-week period. Initial investigations by the finance team assumed vendor error. The security operations team, using a SIEM with UEBA rules, identified that refunds were initiated from an admin account outside business hours, from an unrecognized IP address and with altered browser user-agent strings. The account had been compromised through credential stuffing. Because monitoring detected the anomaly within 72 hours, the company reversed most refunds, revoked the account, notified affected customers and filed a CERT-In report. Post-incident, the platform implemented MFA for all admin accounts, geo-velocity alerts, and automated deactivation of accounts with impossible-travel signals.

Lessons:

  • Centralize logs from all critical systems and retain them for at least 180 days.
  • Build detection rules for India-specific fraud patterns (UPI, wallet refunds, coupon abuse).
  • Test alerting paths quarterly and measure mean time to detect (MTTD).
  • Document monitoring coverage as evidence for RBI/SEBI and ISO 27001 auditors.

💡 Need help implementing this? Contact Singahi for a 20-minute readiness call. We build monitoring programs in 8 weeks.

How we can help

Working toward this?

If a certification or a customer's security questionnaire is what brought you here, tell us where you are. We'll give you an honest read on the work and the timeline, with no obligation.

What happens next

  1. Tell us the trigger

    A questionnaire, an audit date or an investor ask. The short form or a call both work.

  2. A practitioner replies

    A senior practitioner, not a bot, within four business hours.

  3. You get a scoped next step

    An honest view of what the work involves. No pressure, no theatre.