# Private DevOps > Senior DevOps consultancy for startups and scale-ups: Kubernetes, cloud infrastructure, SRE, security, and CI/CD. Direct access to senior engineers, no middlemen and no offshoring, with AI-augmented delivery. Private DevOps LTD is an EU VAT-registered company that has been running production infrastructure since 2017, with engineers who bring 20+ years of hands-on experience. Every client works directly with a senior engineer who owns their stack, not a ticket queue or a rotating bench. We serve clients worldwide in English. Emergency technical support is available 24/7. Primary expertise: Kubernetes, AWS and multi-cloud, Site Reliability Engineering (SRE), security hardening and CVE response, CI/CD pipelines, server management, and e-commerce performance (Magento, WordPress). We offer flexible services and engagement, with no long-term contract required. We work three ways: a monthly DevOps-as-a-Service retainer (hour packages), a scoped project, or a single one-off task or problem. You do not have to sign a retainer to work with us. For a one-off task, tell us what you need and you get a written action plan with a clear time estimate and a fixed price upfront, before any work begins; you approve it, then we do the task. Retainers only make sense once the work is ongoing, they make it cheaper per hour and faster, but they are not a requirement. ## Services - Cloud, Kubernetes & Reliability - [Kubernetes Management](https://privatedevops.com/services/kubernetes-management): Full lifecycle K8s, setup, upgrades, scaling, monitoring, and incident response - [Kubernetes for Next.js](https://privatedevops.com/services/kubernetes-for-nextjs): Production Kubernetes tuned for Next.js, SSR, ISR, and streaming AI endpoints - [SRE Services](https://privatedevops.com/services/sre): Site Reliability Engineering, SLOs and error budgets, production-readiness reviews, incident response, and blameless postmortems - [DevOps as a Service](https://privatedevops.com/services/devops-as-a-service): Senior DevOps on retainer through monthly hour packages - [Infrastructure Management](https://privatedevops.com/services/infrastructure-management): Ongoing ops and maintenance for Magento, Laravel, and legacy systems - [Monitoring & Observability](https://privatedevops.com/services/monitoring-observability): Prometheus, Grafana, Datadog, and ELK, with dashboards, alerting, and APM - [Security & Compliance](https://privatedevops.com/services/security-compliance): Server hardening, CVE triage, zero-trust architecture, and SOC2 readiness - [Disaster Recovery & Backup](https://privatedevops.com/services/disaster-recovery-backup): Tested DR plans, cross-cloud quorum, and restore drills - [Cloud Management](https://privatedevops.com/services/cloud-management): Multi-cloud and hybrid environment management - [AWS Cloud Management](https://privatedevops.com/services/aws-cloud-management): AWS-native operations, compute, databases, CDN, and IAM - [Cloud Cost Optimization](https://privatedevops.com/services/cloud-cost-optimization): Right-sizing, reserved capacity, and architecture changes that cut cloud spend ## Services - Setup, Architecture & Delivery - [Infrastructure Setup](https://privatedevops.com/services/infrastructure-setup): Production-grade cloud infrastructure from scratch using infrastructure as code - [Architecture & Planning](https://privatedevops.com/services/architecture-planning): Infrastructure audits, tech-debt assessment, and migration roadmaps - [CI/CD Pipeline Setup](https://privatedevops.com/services/ci-cd-pipeline-setup): GitHub Actions, GitLab CI, and ArgoCD GitOps workflows - [Migration Services](https://privatedevops.com/services/migration-services): Zero-downtime migrations, on-prem to cloud, EC2 to EKS - [API-as-a-Service on AWS](https://privatedevops.com/services/api-as-a-service): Managed API platforms built on AWS ## Services - Servers & Performance - [Servers Management](https://privatedevops.com/services/servers-management): Patching, monitoring, and security for bare-metal and VPS - [Plesk Servers Management](https://privatedevops.com/services/plesk-servers-management): Managed Plesk hosting operations - [cPanel Servers Management](https://privatedevops.com/services/cpanel-servers-management): Managed cPanel hosting operations - [Server Setup & Optimization](https://privatedevops.com/services/server-optimization): Performance tuning for web servers, PHP, databases, and caching - [Magento 2 Speed Optimization](https://privatedevops.com/services/magento-2-speed-optimization): LCP and Core Web Vitals tuning for Magento 2 stores - [WordPress Speed Optimization](https://privatedevops.com/services/wordpress-speed-optimization): Performance and caching tuning for WordPress ## Industries Who we work with, and what the infrastructure problem looks like for each: https://privatedevops.com/industries - [DevOps for Startups & Scale-ups](https://privatedevops.com/industries/startups): For funded startups from seed to Series C whose engineers are losing half a week to the cloud. Lays the work out on the arc a founder recognises: at pre-seed to seed the risk is not the bill but that one laptop or one leaver takes production with it, so the environment goes into code with a deploy anyone can run and a restore that has actually been tested; from seed to Series A uptime becomes a promise made in writing and the bill starts appearing on the board deck, so monitoring, CI/CD and a cost pass come off the engineers; past Series A due diligence wants to know whether the whole thing rebuilds from code, who has access to what, and where customer data sits. No retainer is required, a single task gets a plan and a fixed price before any work starts, and the page says plainly that a full-time hire usually wins past the first twelve to eighteen months. - [DevOps for SaaS & Software Platforms](https://privatedevops.com/industries/saas): For B2B and B2C SaaS with multi-tenant data, paying customers and uptime written into contracts. Covers the infrastructure work behind a platform whose customers notice an outage before you do. - [DevOps for E-commerce Platforms](https://privatedevops.com/industries/ecommerce): For online retailers and storefronts on Magento 2, WooCommerce and headless stacks, where a slow checkout is a revenue number rather than a performance metric. - [DevOps for Hosting Providers & Agencies](https://privatedevops.com/industries/hosting-providers): For hosting companies and agencies running shared, VPS and managed fleets, where the infrastructure is the product and every incident is somebody else's business too. ## Case Studies Real engagements with measurable outcomes: https://privatedevops.com/case-studies - [Headless Commerce Migration](https://privatedevops.com/case-studies/headless-commerce-migration): Re-platforming an e-commerce store to a headless architecture - [Cloud Security Hardening](https://privatedevops.com/case-studies/cloud-security-hardening): Closing security gaps across a cloud estate - [Infrastructure Cost Optimization](https://privatedevops.com/case-studies/infrastructure-cost-optimization): Cutting cloud spend without losing performance - [The 3 Second Timeout](https://privatedevops.com/case-studies/the-3-second-timeout): Diagnosing a production latency and timeout incident ## Partners Vendor partner programmes: https://privatedevops.com/partners - [Partners - The Programmes We Are In](https://privatedevops.com/partners): The vendor partner programmes Private DevOps LTD belongs to, each stated at the stage it is actually at rather than flattened into one word, and each linked to the vendor's own public register, which both now have, so a reader can check the claim instead of believing it. Carries an explicit note that partner status does not imply endorsement and that recommendations follow the workload rather than the badge, since the same team works across AWS, GCP, DigitalOcean and Hetzner. - [GitLab Partner](https://privatedevops.com/partners/gitlab): Private DevOps LTD is listed in GitLab's Authorized Partner Locator on the Open track, a public entry the page links to directly. Sets out why the platform was chosen rather than why the badge was awarded: repository, pipelines, container registry, package registry, issues and releases arrive as one product that already knows about itself instead of four stitched with webhooks; self-managed GitLab is a first-class option rather than a legacy tier, which is what matters to a team with a regulator or a board that wants the source somewhere it controls; and GitLab CI is YAML the inheriting engineers can still read a year later. The work behind it is migrations onto GitLab including the CI rewrite no importer performs, self-managed GitLab on Kubernetes with upgrades, runners, backups and restore drills, and licensing questions answered before the purchase rather than after. - [DigitalOcean Partner](https://privatedevops.com/partners/digitalocean): Private DevOps LTD is a member of the DigitalOcean Partner Pod and is listed in DigitalOcean's public partner directory at https://www.digitalocean.com/partners/directory/privatedevops, with the partner portal, migration support and opportunity registration. Explains the choice in terms of what a small team can actually operate: a catalogue short enough to hold in your head, where the missing services are mostly the ones a young company would not touch for two years; a DOKS managed control plane, so upgrades, etcd and the API server stop consuming engineer evenings; and documentation written for someone meeting the concept for the first time. The work behind it is DOKS clusters run as production, moves between DigitalOcean and the larger clouds in both directions, and cost and architecture reviews for teams deciding which platform fits their stage. ## Articles (Guides & Tutorials) In-depth technical guides: https://privatedevops.com/articles - [A cPanel Account With Email Access Can Reach Root, and Every Supported Version Is Affected](https://privatedevops.com/articles/cpanel-cve-2026-67401-account-holder-to-root): CVE-2026-67401, published by cPanel on 8 September 2026, is a SQL injection in cPanel's EmailTrack functionality that lets an authenticated account holder with mail-related privileges create arbitrary files on the server, leading to code execution as root and full control of the machine. cPanel lists all supported versions of cPanel and WHM as affected and gives the patched builds as v11.110.0.143, v11.134.0.55, v11.136.0.39, v11.138.0.4 and WP Squared v11.138.1.9. The argument is that the boundary this crosses is the one shared hosting is sold on, since a cPanel account holder is a customer and customer accounts are not supposed to reach root, so for a provider this is a flaw in the product rather than in a tool. Explains why 'authenticated' is weaker than it sounds, because an attacker can simply buy a hosting account, mail privileges ship with practically every package, and compromised customer accounts from credential stuffing supply attackers the provider never chose. Carries the version check with /usr/local/cpanel/cpanel -V, the standard /scripts/upcp update run, and the advice to verify the version rather than trust the automatic update setting. Reads the four absences in the advisory honestly: no CVSS because cPanel do not publish one, no report of exploitation in the wild and a responsible disclosure credit, which unlike the Magento zero-day the week before means this was patched before anyone was using it, no mechanism detail while operators patch, and no MITRE or NVD record yet, which was checked and is normal timing rather than a red flag. - [The Magento Hotfix Is Out, and Adobe Wants Your Payment Gateway Keys Rotated Too](https://privatedevops.com/articles/magento-cve-2026-75650-hotfix-and-key-rotation): Adobe published APSB26-146 on 7 September 2026, a day ahead of its scheduled release, fixing CVE-2026-75650 at CVSS 10.0 with vector CVSS:3.1/AV:N/AC:L/PR:N/UI:N/S:C/C:H/I:H/A:H, Priority 1, CWE-1336 template engine injection, no authentication required, and Adobe states it is aware of exploitation in the wild. Affected are Adobe Commerce 2.4.4 through 2.4.9, Adobe Commerce B2B 1.3.3 through 1.5.3 and Magento Open Source 2.4.6 through 2.4.9 at the August 2026 patch level and earlier. Carries a step-by-step tutorial for applying it, since the fix is a composer hotfix called VULN-39341 rather than a version bump and there is no patch release to upgrade to. The hotfix is published at https://repo.magento.com/patch/VULN-39341-composer-patches.zip and the file inside is VULN-39341_Hotfix_COMPOSER.patch. Establish the strip level first with patch -p1 --dry-run, falling back to -p2 as Adobe documents. On-premises Adobe Commerce and Magento Open Source apply it from the installation root with patch -p1 < VULN-39341_Hotfix_COMPOSER.patch followed by a cache flush, with maintenance mode around it. Adobe Commerce on cloud infrastructure cannot be patched by hand because the next deploy overwrites it, so the file goes into an m2-hotfixes directory in the project root and is committed and pushed, and the build applies it. Confirm with composer require magento/quality-patches then vendor/bin/magento-patches -n status expecting VULN-39341 to read Applied, since a composer patch can fail quietly. Adobe tested the hotfix only against the August 2026 patch levels and asks for a backup and a staging run first. The part the coverage skips is that Adobe's own resolution requires the patch AND an encryption key rotation, and states that rotating the key alone does not invalidate credentials already exposed, so merchants must rotate at the source including payment gateway credentials at Stripe, Braintree, Adyen or PayPal, which is a vendor treating affected stores as potentially compromised rather than merely vulnerable. Explains why the thirteen step order matters (patch before rotating, maintenance mode and cron disabled so nothing re-encrypts mid-rotation), that neither patching nor rotation removes an implant left by an earlier breach, and gives a triage order for merchants who cannot take a maintenance window today. - [Who Holds REPLICATION on Your Postgres, and Why CVE-2026-6471 Makes It Matter](https://privatedevops.com/articles/who-holds-replication-on-your-postgres): CVE-2026-6471 is missing authorization in PostgreSQL logical decoding that lets a non-superuser holding REPLICATION dlopen any file visible to the operating system account running the server, which runs arbitrary code as that account. CVSS 7.2 with vector AV:N/AC:L/PR:H/UI:N/S:U/C:H/I:H/A:H, and the PR:H is the whole point: the attacker needs an account you already created and granted. Corrects three things the coverage gets wrong. It is not a zero-day, since the fix shipped on 13 August 2026 in 18.6, 17.11, 16.15, 15.19 and 14.24; the twelve years describe when logical decoding landed in 9.4 rather than any change in exposure; and the widely repeated claim that a reachable SMB port 445 is a prerequisite is wrong, that being one Windows path to make the library visible over UNC while the advisory says any file visible to the server's OS account. The argument is that high privileges are not a reason to relax, because REPLICATION is granted casually for read replicas, change-data-capture pipelines, analytics vendors and migration tools and then never revoked, so the flaw upgrades a leaked replication credential from reading the write-ahead log to running commands on the database host. Carries the pg_roles query to find every role holding it, the server_version and wal_level checks with the caveat that a default wal_level of replica is a mitigating factor rather than an all-clear, a pg_hba check for replication lines open to 0.0.0.0/0, why patching does nothing about a credential that already leaked, what a replication role should look like, and ALTER ROLE ... NOREPLICATION as the interim move that removes the precondition and is worth keeping afterwards. - [A Magento Zero-Day Is Being Exploited Now, and the First Victim Was Fully Patched](https://privatedevops.com/articles/magento-stylesmuggler-zero-day-under-active-attack): StyleSmuggler, disclosed by Sansec on September 5 2026, is an unauthenticated remote code execution flaw affecting all current versions of Magento Open Source and Adobe Commerce including 2.4.9, with exploitation confirmed from 22:20 on September 4. The detail that matters is that the first confirmed victim ran 2.4.6-p15 with the July and August 2026 patches applied and a clean security:patch-status, so being up to date was not a defence. The chain plants PHP in Magento's template system through the styles properties, reaching it by a POST to /graphql, then has Magento execute it while rendering the standard Payment Transaction Failed Reminder email; nobody opens the message and the attack still succeeds when mail delivery fails, so a quiet mailbox proves nothing. A successful attack starts a Rust implant disguised as [kworker/u:8:0] that connects to a command and control server and waits, persisted by a five-minute cron entry, and at disclosure no security vendor other than Sansec recognised it. Written on 6 September while no fix existed, and carries an update noting that Adobe shipped one on 7 September as APSB26-146 and CVE-2026-75650. Kept as the record of that window, since the compromise checks in it apply just as much after patching. Carries the free detection commands, the trade-off of disabling GraphQL against a headless storefront, why the query-string filter is the fallback, and a caution that the discovering vendor's own advice leads with its two paid products before the free measure. - [What Your Startup Runs, and What You Are Paying For](https://privatedevops.com/articles/what-your-startup-runs-and-what-you-are-paying-for): Written for early startups choosing where to run Kubernetes, and it concedes the AWS case first: more managed services, autoscaling at every layer, the deepest permission model, the widest compliance coverage, and a price that is honest for what it is. The argument is about timing rather than quality. Sets out the structural differences that survive any price change: EKS bills the managed control plane per cluster while DOKS includes it and sells HA as an add-on; AWS meters egress by the gigabyte against a small account-wide allowance while DigitalOcean attaches a transfer allowance to each node and pools it across the cluster, so one bill tracks traffic from month one and the other does not move until the pool is spent; and an EKS cluster's upgrade policy defaults to EXTENDED, so fourteen months after a Kubernetes version leaves standard support the cluster auto-enrols into extended support and the per-cluster rate multiplies for the same work, the alternative being STANDARD and an upgrade on AWS's schedule. The section most coverage omits is the published uptime commitments read side by side: the managed control-plane figures are identical at 99.95% including the credit tiers, but on a single machine DigitalOcean commits to 99.99% per Droplet against 99.5% per EC2 instance, roughly four minutes a month against three and a half hours, and AWS reaches 99.99% only at region level, which is a commitment about infrastructure spread across availability zones that you architect and pay for. Carries the three caveats that keep it honest: the DigitalOcean control-plane figure requires paid HA, the AWS region-level figure is real once you have built for it, and an SLA is a refund policy rather than a measurement of what actually happened. Closes with a decision rule (list the services the product needs in the next twelve months, not at Series B) and an honest read on lock-in, that plain Kubernetes workloads move but provider-specific wiring does not, and by the time you have enough of it to be stuck you are also the company that should be on AWS. - [Eleven npm Packages Compromised in a 53 Minute Attack That Steals Every Credential Your Build Host Can Reach](https://privatedevops.com/articles/npm-compromised-packages-supply-chain-attack): Practical guide to npm supply chain defense, anchored on the August 4, 2026 Shai-Hulud worm that published malicious versions of eleven caching packages (keyv, flat-cache, file-entry-cache, cacheable-request, cacheable, cache-manager, ecto and four @cacheable scoped packages) inside a 53 minute window, harvesting npm tokens, GitHub PATs and OIDC tokens, AWS credentials including IMDS, Kubernetes service account tokens, Vault tokens, SSH keys and .env files via a preinstall hook. The central finding, verified against the npm registry rather than taken from reporting: ten of the eleven malicious releases were PATCH bumps silently eligible for any caret or tilde range, while keyv, the headline package with 604 million monthly downloads, was a MAJOR bump (5.6.0 to 6.0.0) that no existing range could resolve to, making the famous package the structurally safest one and the obscure transitive dependencies the real vector. Explains why a semver range is a standing grant of trust rather than a version, why npm ci and npm install are not interchangeable and which one belongs in a Dockerfile, how to inventory and allow-list install scripts instead of disabling them blind, why lockfile integrity hashes prove nothing about publish-time legitimacy and npm audit signatures answers the other question, exact detection commands including the three payload SHA-256 hashes and the worm's .claude/settings.json and .vscode/tasks.json persistence markers, and why remediation is credential rotation in dependency order rather than package removal. - [Zero-Downtime Kubernetes Deployments: Complete Guide](https://privatedevops.com/articles/zero-downtime-kubernetes-deployments-complete-guide): Rolling updates, blue-green through a service selector switch, canary by ingress weighting, and PodDisruptionBudgets, with the readiness-probe and maxUnavailable settings each one depends on. - [The GitHub to GitLab Migration Nobody Warns You About](https://privatedevops.com/articles/github-to-gitlab-migration-what-moves-and-what-you-rewrite): What GitLab's GitHub importer actually brings across (branches, LFS, issues, pull requests with their reviews and suggestions, labels, milestones, branch protection, collaborators with mapped roles, wikis) and the one thing it does not touch at all: GitHub Actions workflows. Covers the concept mapping GitLab publishes, why two of its seven rows are marked not applicable, the token scopes and prerequisites, the traps (Actions secrets recreated by hand as CI/CD variables, pre-2017 diff comment threading, outside collaborators, status-check rules, SAML attachments), what GitLab Duo and the free agent skill actually do and why neither is a converter, the 76 hours GitLab took to import the Kubernetes repository, and an honest read on when the move is and is not worth making. - [Infrastructure as Code, Terraform vs Pulumi](https://privatedevops.com/articles/infrastructure-as-code-terraform-vs-pulumi): Choosing an IaC tool and the trade-offs - [How to Build a Backup You Have Actually Restored](https://privatedevops.com/articles/backup-you-have-actually-restored): A scheduled job that restores last night's copy into a scratch container, runs a real query against it and alerts when the row count is wrong, plus the failure modes that only surface on restore day (roles and extensions missing from pg_dump, large objects left behind, the encryption key stored next to the backup, a replica that stopped replicating). Commands for pgBackRest and restic. - [How to Set Up SPF, DKIM and DMARC So Mail Actually Lands](https://privatedevops.com/articles/spf-dkim-dmarc-so-mail-actually-lands): Why most setups have all three records and still fail DMARC, explained through alignment: SPF authenticates the envelope sender rather than the From header, so a platform with its own bounce domain passes SPF and aligns with nothing. Covers the ten DNS lookup limit that voids an SPF record, DKIM selector rotation, and the working order from p=none through quarantine to reject. - [How to Find the One Query That Is Actually Killing You](https://privatedevops.com/articles/find-the-query-that-is-killing-you): Ranking by total time rather than duration with pg_stat_statements, because a 3ms query running two million times an hour costs more than a 20 second report running twice. Includes the two defaults that quietly give you less than expected (max 5000, track_planning off), how to read the result as an N+1 or a genuine slow query, and EXPLAIN (ANALYZE, BUFFERS) plus auto_explain. - [Magento 2 Speed Levers That Actually Move LCP](https://privatedevops.com/articles/magento-2-speed-levers-that-actually-move-lcp): The changes that measurably improve Magento LCP - [Hardening a Fresh Ubuntu 24.04 VPS in 15 Minutes](https://privatedevops.com/articles/hardening-fresh-ubuntu-24-04-vps-15-minutes): A fast, practical server hardening checklist - [How to Give Applications AWS Credentials Without Storing Any](https://privatedevops.com/articles/aws-credentials-without-storing-any): Every long-lived access key in your account is a copy waiting to leak, and no amount of rotation discipline fixes that. The alternative is to have no key at all, because each place an application normally needs credentials already has a mechanism that hands it fresh ones on demand. This walks through instance profiles on EC2, task roles on ECS, EKS Pod Identity and IRSA on Kubernetes, and OIDC federation for a CI pipeline, with the trust policy shape for each. It also covers the one condition in the CI trust policy that decides whether the whole thing is secure or theatre. - [How to Reach a Private RDS Without a Bastion Host](https://privatedevops.com/articles/reach-a-private-rds-without-a-bastion): A jump host with a public IP and an open SSH port is the most commonly attacked thing in a lot of AWS accounts, and it exists only so somebody can occasionally run a query. Systems Manager forwards a local port through a managed node to any host that node can reach, so the database stays in its private subnet and nothing accepts inbound connections. This covers the exact command, the agent version and permissions it needs, how it works with no NAT gateway at all, and how to drop the stored database password as well. - [How to Recover an EC2 Instance You Can No Longer SSH Into](https://privatedevops.com/articles/recover-an-ec2-you-cannot-ssh-into): When a box stops answering there is an order to work through, and two of the options only exist if somebody enabled them on a calm afternoon months earlier. This covers what the status checks are telling you, reading console output, the serial console and everything it needs configured in advance, and the volume detach and reattach route as the last resort. The part worth reading before you need it is which mechanisms have prerequisites, because that decides what is available to you at 2am. - [How to Make an S3 Bucket That Cannot Be Deleted by Accident](https://privatedevops.com/articles/s3-bucket-that-cannot-be-deleted-by-accident): Protecting a bucket against a mistake and protecting it against a stolen credential are two different jobs, and the settings that do one do not do the other. This walks through versioning, MFA delete and Object Lock in both of its modes, what each one can and cannot be undone by, and where an attacker with the right permission walks straight through your protection. Several of these settings cannot be reversed once enabled, including one where AWS says the only remaining way to delete the data is to close the account, so the warnings sit next to the commands. - [How to Stop Paying for NAT Gateway Traffic You Do Not Need](https://privatedevops.com/articles/stop-paying-for-nat-gateway-traffic): A large share of NAT gateway spend on a typical account is traffic to AWS services that could have reached those services privately, and it shows up as one anonymous line on the bill. This shows how to tell AWS-bound traffic from internet-bound traffic in Cost Explorer, how to find the exact destination in flow logs, and which endpoint type actually removes the charge. It also covers the traps, including why a bucket in another Region keeps going out through the NAT after you add the endpoint. - [How to Choose Between ALB, NLB and CloudFront for Your Traffic](https://privatedevops.com/articles/choose-between-alb-nlb-and-cloudfront): The three services sit at different layers, accept different protocols, and a handful of the choices you make when you create them cannot be changed afterwards. This is what each one is actually for, where the protocol list makes the decision for you, the two cases where the right answer is a pair of them working together, and the settings that mean rebuilding rather than editing if you get them wrong. - [How to Run Multi AZ So It Actually Survives an AZ Failure](https://privatedevops.com/articles/multi-az-that-actually-survives-an-az-failure): Most AWS accounts are multi AZ on paper already, and then a zone has a bad day and the site goes down anyway. Spreading a deployment across zones and keeping it serving when one disappears are two different properties. This covers what the managed services really do during a zone failure, including which failovers reset every open connection and how long each one takes, the single points that quietly survive a multi AZ design, and the commands to rehearse all of it on purpose. - [How to Migrate a Server to AWS Without a Big Bang Cutover](https://privatedevops.com/articles/migrate-a-server-to-aws-without-a-big-bang-cutover): A big bang cutover is a plan with exactly one attempt in it. The incremental version costs a little more elapsed time and keeps a working rollback available until the very last step. This walks through the inventory that decides whether the cutover is clean, continuous replication that runs while the old server keeps serving, a dress rehearsal you can repeat, the DNS time to live arithmetic you have to do backwards from the cutover date, and the single action that ends the rollback window for good. - [How to Set Up Least Privilege IAM Without Blocking Your Own Team](https://privatedevops.com/articles/least-privilege-iam-without-blocking-your-team): Least privilege earns its reputation for costing a week of tickets whenever someone writes the minimal policy first and discovers what was missing by breaking people's work. The order that avoids that is the reverse. Cap the blast radius, let the team work, collect evidence about what was actually used, and tighten against the evidence. This covers the AWS reporting that supplies the evidence, exactly what data each report is built from and what it silently omits, and the checks that catch an over-tightened policy before it ships. - [How to Upgrade PostgreSQL Major Versions With Almost No Downtime](https://privatedevops.com/articles/postgres-major-version-upgrade-logical-replication): An in-place major upgrade takes your database down for as long as the upgrade runs, and once it has started there is no way back. Logical replication turns that into a cutover you can measure in seconds, with the old server still consistent and still able to take traffic if the first minute goes badly. The method works because the new server is built and caught up while the old one keeps serving. The risk is entirely in what logical replication declines to carry across, so this guide spends most of its time on sequences, DDL, large objects and tables without a replica identity. - [How to Change a Schema on a Busy MySQL Table Without Locking It](https://privatedevops.com/articles/mysql-online-schema-change-busy-table): A plain ALTER on a large InnoDB table can hold up every writer until it finishes, which on a busy table means an outage nobody scheduled. Modern MySQL does far more instantly than most teams realise, so the first job is checking whether you need a tool at all. When you do, the copy-and-swap approach builds a shadow table, keeps it in step from the binary log, and swaps the two at the end. This guide covers what the table has to look like for that to work, and how to stop a migration safely once it is running. - [How to Run Ephemeral CI Runners on Your Own Hardware](https://privatedevops.com/articles/ephemeral-ci-runners-on-your-own-hardware): A build that passes because of something left behind by the previous build is not a passing build, it is a coincidence. Ephemeral runners remove that class of problem by giving every job a machine that has never run anything else. GitHub supports this directly through single-use runner registration and just-in-time configuration, so the runner deregisters itself after one job and your automation disposes of the host. This guide covers both approaches, the Kubernetes version, and the one situation where self-hosted runners are the wrong answer. - [How to Keep a Build Cache That Survives Ephemeral Runners](https://privatedevops.com/articles/build-cache-that-survives-ephemeral-runners): Throwing away the runner after every job is the right call, and it costs you the build cache unless the cache lives somewhere else. On GitHub Actions it already does, which means the real work is writing keys that hit instead of keys that always miss. This guide covers restore-keys and how partial matching actually resolves, the hidden part of a cache key that nobody sets, why a cache saved on a feature branch is invisible to main, and what to do when a bad cache entry starts poisoning every run. - [How to Handle Secrets in CI Without Leaking Them Into Logs](https://privatedevops.com/articles/ci-secrets-without-leaking-them-into-logs): The safest credential in your pipeline is the one that does not exist between jobs. OpenID Connect lets a workflow authenticate directly to a cloud provider and receive a token that expires on its own, which removes the stored key entirely. Masking is the backstop for everything left over, and it is worth knowing exactly where it stops working, because it relies on finding an exact match for the value. This guide covers the short-lived credential setup, the limits of redaction, and what to actually do in the ten minutes after a secret reaches a log. - [How to Restore One Table From a Full Cluster Backup](https://privatedevops.com/articles/restore-one-table-from-a-full-cluster-backup): Somebody emptied one table and the rest of the database is still taking orders. Restoring the whole backup over production would throw away every write since the dump, to fix damage that lives in a single table. This is the side restore instead, with the format choice that decides whether selective restore is even possible, the pg_restore flags that behave differently from the pg_dump ones you know, and the foreign keys and sequences that turn a successful restore into a broken application an hour later. - [What a Real Disaster Recovery Drill Looks Like on a Tuesday](https://privatedevops.com/articles/disaster-recovery-drill-on-a-tuesday): A recovery plan gets tested one of two ways, on a Tuesday morning with a scenario and a stopwatch, or at three in the morning with a customer on the phone. This is the first version, run in working hours and announced in advance, because the failures worth finding are missing documents and expired credentials rather than whether anyone can be woken up. What to scope, the four numbers to measure, what counts as a pass, and the failures a drill turns up nearly every time. - [How to Patch on a Schedule Without a Maintenance Window](https://privatedevops.com/articles/patch-on-a-schedule-without-a-maintenance-window): A quarterly maintenance window means three months of known vulnerabilities waiting for a Saturday night, and then forty machines changing at once while everyone is asleep. Patching continuously and restarting in waves gets fixes on faster and removes the outage entirely. Here is the configuration that actually applies updates rather than only downloading them, how the machine tells you whether a reboot is genuinely required, and how to build waves so no two hosts behind the same load balancer ever go down together. - [How to Read a CVE and Decide in Ten Minutes If It Touches You](https://privatedevops.com/articles/read-a-cve-and-decide-in-ten-minutes): An alarming headline and a severity score are not a decision, and the score cannot become one because the part of it that would describe your environment is the part nobody filled in. This is the ten minute route from a CVE number to a defensible answer, covering what the scoring specification actually says, why your installed version number may be lying about whether you are patched, and which public sources report real exploitation rather than the possibility of it. - [How to Configure Varnish for Magento So It Stops Caching the Wrong Thing](https://privatedevops.com/articles/varnish-for-magento-stop-caching-the-wrong-thing): Varnish in front of Magento serves anonymous pages without touching PHP. The wrong Varnish in front of Magento serves one shopper's cart to everybody, which is worse than having no cache at all. Magento generates its own configuration and most of the work is using that rather than something copied from a forum. This covers what the generated file refuses to cache and why, how invalidation actually works as a ban rather than a purge, the two ways it silently fails, and how to prove a page came from cache. - [How to Back Up and Restore etcd Before You Need It](https://privatedevops.com/articles/etcd-backup-restore-before-you-need-it): A verified etcd snapshot is the difference between a control plane you rebuild in twenty minutes and a cluster you reassemble from memory over a weekend. This walks through taking a snapshot on a self-managed cluster, proving the file is actually restorable rather than merely present, and bringing a cluster back that no longer comes up. It also covers the commands that moved between binaries in recent etcd releases, because the old ones are the ones everybody copies. - [How to Run Postgres on Kubernetes With Point in Time Recovery](https://privatedevops.com/articles/postgres-on-kubernetes-point-in-time-recovery): A nightly dump cannot answer the question that actually gets asked after an incident, which is to put the database back the way it was at 10.42, just before the migration ran. Continuous archiving can. This is an operator based setup on Kubernetes with base backups and write ahead log shipped to object storage, plus recovery to a named timestamp. It also flags the configuration change that makes most copied CloudNativePG YAML out of date. - [How to Set Requests and Limits From Real Usage Instead of Guesses](https://privatedevops.com/articles/kubernetes-requests-limits-from-real-usage): Copied resource blocks cause two expensive problems at once, pods killed for memory they never used and latency nobody can explain. This shows how to read what your workloads actually consume, how CPU throttling shows up in metrics, and why a CPU limit and a memory limit are completely different decisions. The result is a set of numbers you can defend in a review instead of numbers that were inherited from a tutorial. - [How to Run Cron in Kubernetes So Jobs Never Overlap or Vanish](https://privatedevops.com/articles/kubernetes-cronjobs-that-never-overlap-or-vanish): Scheduled work in Kubernetes fails in two quiet ways, a slow job that starts a second copy of itself and a schedule that silently stops firing. Both are configuration, not luck. This covers the CronJob fields that control concurrency, missed schedules, history retention and failure handling, with the actual defaults, so the nightly billing run is still there in the morning and there is a log to read when it is not. - [How to Do Canary Releases Without a Service Mesh](https://privatedevops.com/articles/canary-releases-without-a-service-mesh): You can send five percent of production traffic at a new version, watch the error rate, and roll back in seconds without installing a service mesh. This walks through replica weighted canaries with a progressive delivery controller, real percentage splitting at the edge, an automated pass or fail check against Prometheus, and the rollback path. It also covers what changed when Kubernetes retired Ingress NGINX in March 2026. - [How to Keep a WooCommerce Checkout Up on Black Friday](https://privatedevops.com/articles/woocommerce-checkout-up-on-black-friday): Ten times the traffic barely troubles a WooCommerce catalogue and takes the checkout down. The reason is that the catalogue is served from a page cache while the cart, the checkout and the account pages cannot be, so every one of those requests boots WordPress and hits the database. This guide walks through what WooCommerce itself excludes from caching, which cookies and AJAX endpoints have to be routed past the cache, where the cart and the order actually get stored, and how to turn the checkout ceiling into a number you can size instead of a surprise you discover on the day. - [How to Cut LCP With an Image Pipeline Without Touching the Theme](https://privatedevops.com/articles/cut-lcp-with-an-image-pipeline): On most content and commerce pages the Largest Contentful Paint element is a single image, which means one image decides the score. This guide covers the decisions that belong to the delivery layer rather than to a template, from format negotiation on the Accept header to the size actually credited by the metric, the priority the browser assigns to an image before layout, and the one attribute that quietly ruins a hero. Every attribute here is checked against the HTML standard or web.dev, including which images must never be lazy loaded. - [How to Write Alerts That Wake a Human Only When a Human Is Needed](https://privatedevops.com/articles/alerts-that-wake-a-human): An on-call rota fails long before anyone quits, at the moment the team stops reading the pages. This guide covers the three changes that keep that from happening: alerting on what a customer can feel rather than on a cause, replacing threshold pages with burn rate alerts against an objective, and grouping so that one incident produces one notification. It ends with a query that tells you which of your alerts nobody has acted on, and what to do with them. - [How to Run OpenTelemetry on a Small Cluster Without a Vendor Bill](https://privatedevops.com/articles/opentelemetry-small-cluster-no-vendor-bill): A self-hosted OpenTelemetry path is genuinely within reach for a team that does not want its observability metered by somebody else. This guide covers the parts that matter, starting with what the collector actually does with a pipeline, then the difference between running it as an agent and as a gateway, where traces, metrics and logs can land, and how to work out how much disk a retention window needs before you commit to it. Every component name and default here is checked against the OpenTelemetry and backend documentation. - [How to Set SLOs and Error Budgets for a Team of Five](https://privatedevops.com/articles/slos-error-budgets-team-of-five): Reliability targets written for a company of a thousand do not survive contact with a team of five. This guide keeps the arithmetic and drops the ceremony, working through how to pick one or two objectives a customer would actually notice, how to turn a percentage into minutes and into a count of failed requests, and how to use the remaining budget to settle the argument about whether to ship the feature or fix the bug. Every calculation is shown so you can check it against your own numbers. - [How to Run Magento 2 on Kubernetes Without Overpaying](https://privatedevops.com/articles/run-magento-2-on-kubernetes-without-overpaying): Most Magento clusters cost more than they need to because every tier gets replicated as if every tier were the bottleneck, when only PHP-FPM ever is. This is the shape that keeps the bill honest, with sessions and cache moved to Valkey or Redis, media moved to object storage instead of a shared filesystem, and cron running on exactly one scheduler because Adobe documents that it can only run on one node. It also covers the case for not doing this at all, since a single well-sized server with a warm standby reaches the same uptime for a single steady store. - [How to Tune a Linux Server That Falls Over Only at Peak](https://privatedevops.com/articles/tune-a-linux-server-that-falls-over-only-at-peak): Load average under one, CPU half idle, and for four minutes every evening the site returns errors. The limits that bite at peak are queues and counters rather than utilisation, which is why no dashboard you own is measuring them. This guide shows where to read the accept queue overflow counter, the PHP-FPM log lines that name the pool ceiling in plain words, the ephemeral port range, and the connection tracking table, along with the sysctl and nginx directives that raise each one. Every observation command comes before the change it justifies. - [How to Automate Certificates for Hundreds of Client Domains](https://privatedevops.com/articles/automate-certificates-for-hundreds-of-client-domains): Maximum certificate lifetimes are shrinking on a published schedule, so the number of renewals an agency or host runs each year is about to multiply twice over. This guide covers the pattern that survives it: DNS-01 validation with CNAME delegation, so a customer makes one record once and you never hold their zone credentials. It also covers which Let's Encrypt rate limit you meet first depending on how your names are shaped, why the authorization failure limit turns a small mistake into an outage, and how ACME Renewal Information takes the renewal schedule out of your hands entirely. - [How to Move Off cPanel Without Breaking Mail](https://privatedevops.com/articles/move-off-cpanel-without-breaking-mail): A web request that lands on the old server serves a slightly stale page. A message that lands on the old server sits in a mailbox the customer can no longer reach, and it never moves on its own. This is the migration order that keeps every message reachable throughout, with the inventory items that break silently, the TTL work that has to happen days ahead, repeated incremental imapsync passes across the cutover, and the authentication records that must travel with the mail. The DKIM private key does not come with the mailboxes, and that is the step most migrations discover afterwards. - [How to Block an Attack Without Blocking Your Own Client](https://privatedevops.com/articles/block-an-attack-without-blocking-your-own-client): The rate limit that finally stops the credential stuffing is also the one that locks out your customer's head office on Monday morning, or quietly drops a payment provider's webhook and leaves a hundred orders unpaid. This guide builds the version that does not do that, starting with getting the real client address right behind a CDN, then an allowlist that works because nginx does not account requests with an empty key, a dry run week that shows you who you were about to break, and a fail2ban jail with the same allowlist repeated. It ends with the list of addresses that must never be banned. - [How to Kill Password SSH Across a Fleet, Once](https://privatedevops.com/articles/kill-password-ssh-across-a-fleet): Password authentication on SSH is the single setting that turns a leaked or guessed credential into a shell. Turning it off is five lines of config; doing it across a fleet without locking yourself out is the part that needs an order. This is that order, including the check that proves keys work before you disable the fallback, why AuthenticationMethods beats PasswordAuthentication alone, and how to leave one deliberate way back in. - [How to Drain a Node Without Dropping a Single Request](https://privatedevops.com/articles/drain-a-node-without-dropping-a-request): kubectl drain politely evicts your pods and your users still see errors, because a pod that is terminating is not automatically a pod that has stopped receiving traffic. The gap between those two states is where the dropped requests live. This is how to close it with a pre-stop delay, a grace period long enough for real requests, and a disruption budget that stops the node from taking your last replica with it. - [How to Pool Connections Without Pgbouncer Lying to You](https://privatedevops.com/articles/pgbouncer-pooling-without-surprises): PgBouncer in front of Postgres is the standard fix for connection exhaustion, and the standard way to break an application in ways that only appear under load. The reason is pool mode: session pooling is safe and pools almost nothing, transaction pooling gives you the numbers you wanted and quietly removes session state your ORM assumed. Here is what each mode actually does, which features stop working, and the settings that decide whether it helps. - [How to Set Up Budgets and Anomaly Detection Before the Bill Surprises You](https://privatedevops.com/articles/budgets-and-anomaly-detection-before-the-bill-surprises-you): Finding out about a spend problem from the invoice means finding out weeks late. AWS gives you two different mechanisms for catching it earlier, and they are not interchangeable. A budget fires when a number you chose is crossed, while Cost Anomaly Detection models what your spend normally looks like and tells you when the shape changes. Covers setting up both from the CLI, the delay each one carries between the spend happening and the alert arriving, who should be on which notification, and the charges neither one will catch for you. ## News (Security & Industry) Latest DevOps and security news, fact-checked against primary sources: https://privatedevops.com/news - [Hetzner IP Space Hijacked to Poison a Virtualizor Update](https://privatedevops.com/news/hetzner-ip-space-hijacked-virtualizor-update-poisoned): Between 28 and 30 August 2026, AS62390 (NexonHost) announced 162.55.80.0/24, part of Hetzner's address space carrying Softaculous update and billing systems, via transit provider AS6204 (Zet.net), diverting traffic to an attacker-controlled server across two waves separated by an eleven-hour lull. The attacker obtained a genuine Let's Encrypt certificate because the certificate authority's automated domain validation was routed through the hijack too, so affected clients saw no warning, and a malicious Virtualizor update package reached a small number of installations because update clients did not cryptographically verify packages. The indicator of compromise is /etc/systemd/system/java-jre-update.service. The post's original contribution, which neither vendor advisory covers (both mention RPKI, ROA and ROV zero times), is why origin validation did not stop it: the attacker kept AS24940 (Hetzner) as the apparent origin, confirmed independently against RIPE routing history which returns exactly one origin network across the whole incident, and per RIPEstat's historical record Hetzner's signed record for 162.55.0.0/16 permitted subdivision down to /24 and had done since March 2021, so the hijacked announcement satisfied both the origin and the prefix-length test rather than slipping past them. Hetzner has since capped the record at /16, verified against two independent sources on 2 September, which also forecloses the more-specific /24 announcement Hetzner itself used as mitigation during the incident. Corrects the common framing that Hetzner was breached, gives the honest diversion figures (about 72 percent of RIPE's 368 collector peers at peak while a wave was active, 28 percent time-weighted across the full 33 hours), separates what Softaculous document from what Hetzner has said publicly (nothing on their status page or in their pressroom), and covers remediation. - [OVHcloud Prices Rise on 1 October, at Your Renewal Date](https://privatedevops.com/news/ovhcloud-price-increase-october-2026-renewal): What OVHcloud actually announced, read from their own pricing tables rather than the coverage. Server configurations rise between 5 and 158 percent depending on range and generation, while memory options on some ranges rise between 162 and 652 percent, so the widely quoted 87 percent is neither the ceiling nor the right figure to plan with. Covers the three dates (1 July for memory and storage options, 11 August for new configurations, 1 October for servers already running, applied at each subscription's renewal), which ranges are untouched (Kimsufi, So you Start, Rise and everything before Gen 2024), why DDR5 costs six times what it did a year ago, and what to check before October. - [cPanel and Plesk Both Patched a Root Escalation the Same Day](https://privatedevops.com/news/cpanel-plesk-root-escalation-august-2026): On August 27, 2026 cPanel fixed CVE-2026-65643 and Plesk fixed CVE-2026-67394, both letting a customer account reach root on a shared server. The cPanel bug needs only an account that can add a parked or addon domain and affects all supported versions; the Plesk one needs shell access, does not affect Plesk for Windows, and can be blocked by disabling shell without patching. - [Your CI History Now Expires With Your Build Artifacts](https://privatedevops.com/news/github-actions-retention-checks-workflow-runs): From October 1, 2026 GitHub applies the Actions retention setting to checks, workflow runs and statuses, which previously survived 400+ days regardless of configuration. Teams that lowered retention to control artifact storage cost will lose build history at the same interval unless they raise the number or export what they need. - [GitHub Now Revokes Stolen Tokens Without Locking Out the Rest](https://privatedevops.com/news/github-credential-revocation-by-token-type): On August 18, 2026 GitHub made credential revocation and deauthorization by token type generally available. Previously the credential kill switch applied to all of a user's credentials at once, so containing one leaked token also took their SSH keys and OAuth authorizations. Enterprise owners, organization admins and holders of the Manage enterprise credentials permission can now act on one type at a time: personal access tokens, SSH keys, OAuth app tokens or GitHub App user access tokens. The post separates the two capabilities most coverage flattens: deauthorization removes SSO authorizations for a credential type while the credential still exists (reversible, the right first move), and revocation deletes the user-level credentials themselves (GitHub's example is deleting every personal access token for one Enterprise Managed User without touching their SSH keys). Both work from the web UI and the REST APIs. The quieter half of the release is organization-level parity: every bulk revocation action that was enterprise-only now exists at organization level too, which widens who this helps to single-organization teams. All actions land in the audit log and affected users are emailed. Ties the capability to the August 4 npm supply chain worm, which harvested exactly these token types, and closes with two cheap preparations: check who actually holds the permission, and put the API call rather than a UI walkthrough in the runbook. - [GitHub Now Migrates GitLab Repositories in One Command](https://privatedevops.com/news/gitlab-to-github-migration-generally-available): On August 3, 2026 GitHub made self-service GitLab migrations generally available via GitHub Enterprise Importer and the `gh gl2gh` CLI extension, replacing the previous Expert Services-only path. Sources are gitlab.com and currently maintained non-EOL GitLab Self-Managed versions; archives stage in GitHub-owned blob storage or your own S3/Azure account. What migrates: Git source and history, wiki, commit comments, issues with state and milestone events, merge requests converted to pull requests with reviewers, approvers, timeline events and reactions, releases and assets, attachments, and members as mannequins. The post leads with what does not: GitHub's docs state that .gitlab-ci.yml has no automatic GitHub Actions equivalent, so pipelines, schedules, CI/CD variables, triggers, job traces, artifacts and webhooks are all excluded, which means the pipeline rewrite is the migration rather than a follow-up task. Second decisive limitation: the destination must be GitHub Enterprise Cloud because migrations to GitHub Enterprise Server are not supported, which rules out the residency and air-gap cases that put teams on self-managed GitLab in the first place. Also covers lossy conversions (threaded discussions flattened, review comments downgraded when diff data is absent, only the latest diff exported), Git LFS pointer files arriving without their binaries, and why merge request count rather than repository count sets the timeline, with the `gh gl2gh inventory-report` command and the hard limits (2 GiB commit and push, 255 byte refs, 100 MiB files, 40 GiB archive). - [The New cPanel Critical Bug Needs a Valid Login and Still Outscores April's Unauthenticated Root Flaw](https://privatedevops.com/news/cpanel-cve-2026-58048-database-privilege-escalation): CVE-2026-58048, published July 31, 2026 by HackerOne on behalf of WebPros, is a 9.4 critical privilege escalation in cPanel and WHM plus WP Squared, reported by Vincent55 Yang. Improper preservation of SQL mode when renaming a database causes SQL to execute in root context (CWE-89), so an account holder with the MySQL or MariaDB feature can read, alter or drop any database on the server, extract credentials and plant persistent database objects. The post's core finding is a comparison no other coverage makes: this authenticated bug scores 9.4 while April's unauthenticated CVE-2026-41940 bypass scores 9.3 on the same CVSS 4.0 scale, and the entire difference sits in the subsequent-system metrics (SC:H/SI:H/SA:H versus SC:N/SI:N/SA:N), meaning the CNA scored this one as leaving the account and taking the host with it. Explains why "requires authentication" is a price list rather than a barrier on shared hosting and reseller servers, lists the first fixed build per release tier (11.110.0.137, 11.118.0.71, 11.126.0.78, 11.134.0.48, 11.136.0.32, WP Squared 11.138.1.6) plus the 137.9999.98 targeted security release of July 29, covers the quieter companion CVE-2026-58047 HTTP request smuggling (5.6, CWE-444), and argues that nightly upcp automatic updates are a statement about configuration while a build number is a statement about state. - [GitHub Ships Stacked Pull Requests and One Limitation Decides Whether Your Team Can Actually Use Them](https://privatedevops.com/news/github-stacked-pull-requests-public-preview): GitHub moved stacked pull requests into public preview on July 30, 2026, rolling out to all repositories over the following days, with merge queue support arriving progressively. A stack is an ordered chain of dependent pull requests in one repository: the bottom targets main, each layer above targets the branch below, each shows only its own diff, and they must merge bottom up. Merging a lower layer automatically rebases and retargets the ones above, which is the manual, error-prone part GitHub has taken over. Driven by `gh extension install github/gh-stack`, and also available on github.com, mobile and to coding agents. The post argues the real benefit is not reviewability but that one objection no longer blocks an entire change, and it leads with the limitation the announcement omits: all branches must be in the same repository, so cross-fork stacks are not supported and the fork-based contribution model is ruled out entirely. Also covers the GitHub Desktop gap, the new merge API needed for API-driven merges, and what teams on GitLab or Graphite can and cannot replicate. - [On July 30 AWS Quietly Trims a Dozen Services and Walking Back Its Own AI Bets Is the Real Story](https://privatedevops.com/news/aws-service-retirements-july-2026): On June 30, 2026 AWS closed roughly a dozen services to new customers effective July 30, mostly its own first-generation AI and search products (Amazon Kendra, Amazon Q Business, Amazon Bedrock Agents renamed Classic, plus many SageMaker AI features). Retire means three different things here: maintenance mode (closed to new customers, existing keep running fully supported, no new features), sunset (a real end-of-support date, migrate), and already ended June 30. The post sorts every service into its bucket, argues the AI cull is a consolidation onto Bedrock rather than a retreat, and details the two real migrations: WorkSpaces PCoIP to Amazon DCV by October 31, 2027, and Kendra to Bedrock Managed Knowledge Bases as an assessment rather than a drop-in. - [The Node.js 20 Lambda Deadline Everyone Is Citing Is Wrong and the Real Risk Already Started](https://privatedevops.com/news/aws-lambda-node-js-20-deprecation-block-dates-2027): The widely cited AWS Lambda Node.js 20 deadline of August 31 and September 30, 2026 is out of date. Per current AWS docs, block-create moved to February 1, 2027 and block-update to March 3, 2027. The date that actually matters is April 30, 2026, when the Node.js 20 language runtime stopped receiving security patches, so functions are already running unpatched. Covers what each of the four events (deprecation, block create, block update, invocation) really means, the OS-versus-language-runtime patching nuance most guides miss, a CLI sweep to find every affected function across regions, and the safe migration to nodejs22.x or nodejs24.x. - [The wp2shell WordPress RCE Is Real, but Three Conditions Decide Whether Your Site Is Actually Exposed](https://privatedevops.com/news/wp2shell-wordpress-core-rce-cve-2026-63030-who-is-exposed): wp2shell (CVE-2026-63030) chains a REST API batch route confusion with the author__not_in SQL injection (CVE-2026-60137) into a pre-auth RCE on WordPress core, patched July 17, 2026 in 6.9.5, 7.0.2 and 6.8.6. An anonymous request can run code on a default install, but three conditions decide real exposure: the version (6.9.0-6.9.4 or 7.0.0-7.0.1 for the RCE, 6.8.x is SQLi-only), whether a persistent object cache is in use (Cloudflare confirmed the RCE path is only reachable without one), and whether WordPress force-pushed auto-updates already patched you. Covers the disputed CVSS scores (WPScan vs CISA-ADP), a two-minute wp-cli check, and post-patch integrity steps. - [nginx Patches Three CVEs, One a 9.2 Critical Bug That Sat Hidden in the Code Since 2011](https://privatedevops.com/news/nginx-cve-2026-42533-critical-buffer-overflow-patched): Three CVEs patched July 15, 2026, in nginx 1.30.4 stable and 1.31.3 mainline. CVE-2026-42533 is a 9.2 critical heap buffer overflow in map regex matching, present since 0.9.6 in 2011 and needing no special module. CVE-2026-60005 is an uninitialized-memory bug triggered by unnamed regex captures through either the slice directive or ordinary background cache updates, which is wider than the slice-module-only framing most coverage uses. CVE-2026-56434 is a use-after-free in the SSI filter, present since 2009, triggered by a crafted proxied backend response. Includes the exact nginx changelog wording, the config patterns to audit for, named-capture mitigation, five read-only audit commands, a fleet audit loop, and the CVSS 3.1 vs 4.0 split. - [A 16-Year-Old KVM Bug Called Januscape Lets a Guest VM Break Out to the Host on Intel and AMD](https://privatedevops.com/news/januscape-cve-2026-53359-kvm-guest-to-host-escape): CVE-2026-53359, a 16-year-old use-after-free in the Linux KVM shadow MMU on Intel and AMD, rated a guest-to-host escape by Canonical. The public proof of concept crashes the host; full takeover is claimed but not public. Upstream fixed July 4, distro kernels pending. Who is exposed (nested virtualization plus guest root), the disable-nested-virt mitigation, and the two coupled CVEs to patch. - [Magento Open Source 2.4.6 Goes End of Life on August 11 and There Is No Extended Support to Save You](https://privatedevops.com/news/magento-open-source-2-4-6-end-of-life-august-2026): Magento Open Source 2.4.6 loses all patches on August 11, 2026, and unlike Adobe Commerce it has no paid extended support. What end of support means for security and PCI, where to upgrade (2.4.8 or 2.4.9, not 2.4.7), the PHP jump it forces, and how to plan the move before the deadline. - [A 15-Year-Old Linux Kernel Bug Called GhostLock Gives Any Local User Root and Escapes Containers](https://privatedevops.com/news/ghostlock-cve-2026-43499-linux-kernel-root-who-is-at-risk): CVE-2026-43499, a use-after-free in the kernel rtmutex code (CVSS 7.8 High) that lets any local user reach root. Quietly patched upstream in May but now has a public exploit that Nebula reports escapes containers. Who is actually at risk (multi-tenant, CI, container hosts) and the patch triage. - [The New libssh2 SSH Flaw Is Client-Side, Not Your sshd, and apt upgrade Will Not Fix the Copies That Matter](https://privatedevops.com/news/libssh2-cve-2026-55200-client-ssh-who-is-affected): CVE-2026-55200 is an out-of-bounds write in the libssh2 client library (ssh2_transport_read, no upper bound on packet_length, CWE-680) that a malicious or compromised SSH server can use to corrupt a connecting client and possibly run code; a public PoC is out but there is no confirmed in-the-wild use, severity is disputed (VulnCheck 9.2 critical vs Red Hat 7.1 moderate), the fix is upstream commit 97acf3d with no 1.11.2 release, RHEL 8+ and Ubuntu LTS 20.04 to 24.04 are not affected while Debian bookworm and Alpine 3.19 to 3.20 are, and the real cleanup is the statically linked and vendored copies in curl, Git, PHP, backup agents, and containers that a distribution update never touches - [OpenAI Is Deleting GPT-5 and o3 in December and the Teams That Pinned Their Models Break First](https://privatedevops.com/news/openai-retires-gpt-5-o3-snapshots-december-2026): Six GPT-5 and o3 snapshots leave the first-party OpenAI API on December 11, 2026; pinned dated model IDs break loudly while floating aliases silently change behavior, plus how to inventory, flag, and eval the swap - [Helm 3 Is Going End of Life and Helm 4 Will Quietly Break the Flags Your Pipeline Depends On](https://privatedevops.com/news/helm-3-end-of-life-helm-4-breaking-changes): Helm 3 stops getting bug fixes on September 9, 2026 and security fixes on February 10, 2027; Helm 4 renames --atomic and --force and changes registry login, breaking CI/CD automation not just interactive use - [GitHub Will Stop Running Your Jobs When Your Self-Hosted Runners Fall Behind](https://privatedevops.com/news/github-actions-self-hosted-runner-minimum-version): GitHub Actions enforces a minimum self-hosted runner version (2.329.0) and a 30-day update window; outdated runners stop executing jobs at the July 31 (Data Residency) and September 25, 2026 enforcement dates - [Kubernetes 1.35 Is the Last Version That Runs containerd 1.x and Most Clusters Have Not Noticed](https://privatedevops.com/news/containerd-1-x-eol-kubernetes-1-35-last-version): containerd 1.x reaches end of support and Kubernetes 1.35 is the last release supporting it; containerd 2.0 drops Docker Schema 1 pulls and the CRI v1alpha2 API, so plan the runtime migration before your next cluster upgrade - [Azure Returns 410 Gone for GPT-4o on October 1 and Auto-Upgrade Skips the Deployments That Matter](https://privatedevops.com/news/azure-openai-gpt-4o-retirement-october-2026): Azure OpenAI retires GA gpt-4o and gpt-4o-mini on October 1, 2026; Standard deployments auto-upgrade but Provisioned (PTU) and NoAutoUpgrade deployments must be migrated by hand or return 410 Gone - [Stay on EKS 1.33 and AWS Starts Billing You 6x From August](https://privatedevops.com/news/aws-eks-1-33-extended-support-6x-cost): Amazon EKS 1.33 leaves standard support on July 29, 2026 and auto-enrolls into extended support at roughly six times the control-plane rate; audit deprecated APIs, addons, and the node AMI path before the surcharge - [OpenSSL 3.0 Stops Getting Security Fixes in September and You Probably Still Ship It](https://privatedevops.com/news/openssl-3-0-end-of-life-september-2026): OpenSSL 3.0 reaches upstream end of life on September 7, 2026; distros backport their own builds, so the real exposure is self-compiled, vendored, or container-bundled OpenSSL and language runtimes that ship their own - [From September 11 the EU Cyber Resilience Act Puts a 24-Hour Clock on Exploited Vulnerabilities and Most Makers Have No Runbook](https://privatedevops.com/news/eu-cyber-resilience-act-24-hour-reporting): The first binding CRA obligation starts September 11, 2026; makers of products with digital elements must report actively exploited vulnerabilities in 24 hours and severe incidents, with full notice in 72 hours, or face fines up to 15 million euros or 2.5 percent of turnover - [AWS Shuts Down App Mesh for Good on September 30 and Your Service Mesh Goes With It](https://privatedevops.com/news/aws-app-mesh-shutdown-september-2026): AWS App Mesh reaches full end of support on September 30, 2026; the console and resources stop working, ECS workloads move to Service Connect, and EKS users plan a VPC Lattice or self-managed Envoy migration - [RDS MySQL 8.0 Leaves Standard Support in July and AWS Auto-Enrolls You Into a Paid Bill](https://privatedevops.com/news/aws-rds-mysql-8-0-end-of-standard-support): Amazon RDS for MySQL 8.0 ends standard support on July 31, 2026 and auto-enrolls instances into paid Extended Support billed per vCPU-hour; plan the 8.0 to 8.4 upgrade before August 1 - [AI Agents Broke GitHub and Gave Elon Musk a Shot at Owning the Code You Write](https://privatedevops.com/news/ai-agents-github-musk-software-stack): How the agentic coding boom broke GitHub's pricing and infrastructure, cracked open the Git layer (Cursor Origin), and let SpaceX buy Cursor and tie it to Grok plus Tesla energy and compute. A fact-checked map of the consolidation and what it means for your team (cost control, model and forge lock-in, data governance) - [Cursor Origin Is a Git Forge for AI Agents Worth Watching](https://privatedevops.com/news/cursor-origin-git-forge-for-ai-agents): Calm take on Cursor's June 16, 2026 launch of Origin, a Git-compatible forge built for AI agents committing in parallel, the same-day SpaceX acquisition context, and why teams should watch it but not switch yet (waitlist, GA fall 2026) - [If You Use Gravity SMTP On WordPress Rotate Your Email API Keys Now](https://privatedevops.com/news/gravity-smtp-wordpress-cve-2026-4020-rotate-api-keys): CVE-2026-4020 leaks email provider API keys from the Gravity SMTP plugin to unauthenticated attackers and is being mass-exploited; patch to 2.1.5 and rotate keys - [Two Critical NGINX Bugs Dropped This Week And Who Is Actually At Risk](https://privatedevops.com/news/nginx-critical-cves-2026-42530-42055-who-is-at-risk): Calm triage of CVE-2026-42530 (HTTP/3) and CVE-2026-42055 (HTTP/2 proxy and gRPC), which versions and configs are actually exploitable, and why it is a denial of service for most - [GitHub's July 15 OIDC Change Will Not Break Your Existing AWS Deploys](https://privatedevops.com/news/github-immutable-oidc-sub-claims-aws-deploys): What GitHub's immutable OIDC subject claims actually change, the three triggers that flip you to the new format, and how to future-proof your AWS IAM trust policy - [Hetzner More Than Doubled Some Cloud Prices Today And What To Do About It](https://privatedevops.com/news/hetzner-june-2026-cloud-price-increase-what-to-do): Verified June 15, 2026 Hetzner cloud repricing, which lines jumped most, who is affected, and how to scale around it - [You Can Now Run 200B AI Models On A Desktop Without The Cloud](https://privatedevops.com/news/amd-ryzen-ai-max-395-run-large-models-locally): What AMD's Ryzen AI Max+ 395 with 128GB unified memory really means for local AI, cost, and privacy - [Ansible authorized_key Privilege Escalation, CVE-2026-11837](https://privatedevops.com/news/ansible-authorized-key-privilege-escalation-cve-2026-11837): Local privesc in ansible.posix, who is at risk and how to mitigate - [npm v12 Blocks Install Scripts by Default](https://privatedevops.com/news/npm-v12-blocks-install-scripts-prepare-your-ci): What changes in npm v12 and how to prepare your CI - [Spectra Gutenberg Blocks RCE, CVE-2026-7465](https://privatedevops.com/news/spectra-gutenberg-blocks-rce-cve-2026-7465): WordPress RCE triage by trust level ## Blog (Perspective & How-To) DevOps perspective and how-tos: https://privatedevops.com/blog - [Why So Many Node Apps Still Run on EC2 and PM2, and When It Is Time for ECS or EKS](https://privatedevops.com/ec2-pm2-node-setup-when-to-move-to-ecs-or-eks): Why the old AWS setup (EC2 instances running Node.js under PM2 behind a load balancer) persists for years, the risks it accrues (unpatched OS and runtime, snowflake servers, manual risky deploys, oversold resilience, cost), when it is genuinely fine to keep, and when to move to ECS or EKS, with an incremental no-rewrite migration path. - [Debugging an Airbyte 'All the Defined Primary Keys Are Null' Outage](https://privatedevops.com/airbyte-primary-keys-are-null-outage-postmortem): A two-bug postmortem of an Airbyte MySQL to BigQuery CDC pipeline that would not sync, why the error was never a data problem, the mid-incident connector upgrade we own, and how systematic elimination found the expired cluster certificate and the bad connector version - [Twenty Five Years From Compiling Apache by Hand to Prompting an AI](https://privatedevops.com/twenty-five-years-compiling-apache-to-prompting-ai): How DevOps changed and why judgment still matters in the AI era - [SRE vs DevOps and Why the Difference Decides Your Uptime](https://privatedevops.com/sre-vs-devops-difference-uptime): What SRE actually is and when you need it - [How to Start Doing SRE With SLOs and Error Budgets](https://privatedevops.com/how-to-start-sre-slos-error-budgets): A practical first-steps guide with a worked example ## Free DevOps Tools 22 free, no-signup tools for DevOps and infrastructure teams (speed tests, cost calculators, assessments, SSL, DNS, email deliverability): https://privatedevops.com/tools - [Website Speed Test](https://privatedevops.com/tools/website-speed-test): Enter any URL to measure server response time, TTFB, page size, and compression. Instant results, no signup required. - [WordPress Speed & Security Test](https://privatedevops.com/tools/wordpress-speed-test): Deep scan for WordPress sites - server speed, caching plugin detection, REST API exposure, XML-RPC, login page security, theme, and WooCommerce detection. - [Magento Performance & Cache Test](https://privatedevops.com/tools/magento-speed-test): Deep scan for Magento stores - Full Page Cache status, Varnish detection, admin panel exposure, REST API security, CSS/JS merge, and theme analysis. - [Security Scanner](https://privatedevops.com/tools/website-security-scanner): Scan your website for security headers, SSL configuration, and common vulnerabilities. Instant security grade. - [Server Performance Test](https://privatedevops.com/tools/server-performance-test): Full server health check - response time, TTFB, page weight, compression, SSL grade, security headers, and HTTP version. Similar to Pingdom but free. - [Plesk Server Scanner](https://privatedevops.com/tools/plesk-server-scanner): Deep scan for Plesk servers - panel detection on port 8443, Plesk version, web server type, PHP version, webmail, WAF/ModSecurity status, and performance metrics. - [cPanel Server Scanner](https://privatedevops.com/tools/cpanel-server-scanner): Deep scan for cPanel servers - panel on port 2083, WHM access, cPanel version, webmail, PHP version, web server type, ModSecurity, and AutoSSL status. - [K8s Cost Calculator](https://privatedevops.com/tools/kubernetes-cost-calculator): Estimate your Kubernetes infrastructure cost and see how much you could save with managed K8s. - [Downtime Cost Calculator](https://privatedevops.com/tools/downtime-cost-calculator): Calculate how much downtime is really costing your business. See the impact of proactive monitoring. - [RTO/RPO Calculator](https://privatedevops.com/tools/disaster-recovery-calculator): Estimate how long it would take to recover your system after a failure, and how much data you could lose. Based on your backup strategy. - [Monitoring Readiness](https://privatedevops.com/tools/monitoring-readiness-assessment): Check how well your infrastructure is monitored. Answer 8 quick questions to find blind spots. - [Migration Planner](https://privatedevops.com/tools/cloud-migration-planner): Get an estimated timeline and risk assessment for your cloud migration in under a minute. - [Cloud Health Assessment](https://privatedevops.com/tools/cloud-health-assessment): Quick 6-question assessment of your cloud environment health. Find gaps before they become incidents. - [AWS Cost Estimator](https://privatedevops.com/tools/aws-cost-estimator): Estimate your monthly AWS bill based on your resource usage. Get optimization recommendations. - [SSL Certificate Checker](https://privatedevops.com/tools/ssl-certificate-checker): Check any domain's SSL certificate - expiry date, certificate chain, trust status, SANs, and cipher details. Instant results. - [CSR Decoder](https://privatedevops.com/tools/csr-decoder): Paste your Certificate Signing Request (CSR) to decode and verify its contents - Common Name, Organization, Key Size, and more. - [Certificate Decoder](https://privatedevops.com/tools/ssl-certificate-decoder): Paste an SSL certificate in PEM format to decode its details - subject, issuer, validity dates, SANs, key info, and fingerprints. - [Certificate Key Matcher](https://privatedevops.com/tools/certificate-key-matcher): Verify that an SSL certificate and private key are a matching pair. Paste both to check compatibility before installation. - [Mail Score Checker](https://privatedevops.com/tools/email-configuration-checker): Check your domain's email setup - MX records, SPF, DKIM, DMARC, MTA-STS, BIMI, and reverse DNS. Get a mail deliverability score with actionable recommendations. - [Email Blacklist Checker](https://privatedevops.com/tools/email-blacklist-checker): Check if your domain or IP is blacklisted on 25+ major email blacklists (DNSBL). Instant results with clear listed/clean/timeout status. - [DNS Record Lookup](https://privatedevops.com/tools/dns-record-lookup): Look up all DNS records for any domain - A, AAAA, MX, TXT, NS, CNAME, and SOA records. Detect DNS provider and view TTL values. - [Reverse DNS Lookup](https://privatedevops.com/tools/reverse-dns-lookup): Check reverse DNS (PTR) records for any IP address or domain. Verify that your mail server IP has proper PTR configuration for email deliverability. ## Contact - Website: https://privatedevops.com/contact - Post a task: https://privatedevops.com/request - one page for posting a piece of work. Pick what it is on (a server, a website, cloud infrastructure, or something else) with optional tags such as MySQL, PHP, Plesk, cPanel, SSL, WordPress, Magento, AWS, Kubernetes or Terraform, give it a title and describe it. A website request also asks for the URL and a cloud request asks which cloud. Then say how soon it is needed (flexible at the standard rate, within a week, or urgent at a priority rate), whether you want a one-off task, hourly help or ongoing monthly support, and whether to reply by email only or also on Slack or Teams. Finally say whether you are an individual or a company; a company adds a country of registration and a VAT or company number so an invoice can follow. You get back a scope, an approach and a price, usually within one working day. No credentials are ever requested through the form. - Email: hello@privatedevops.com - Emergency technical support (24/7): https://privatedevops.com/emergency-technical-support - Company LinkedIn: https://www.linkedin.com/company/privatedevops ## Optional - [About](https://privatedevops.com/about): Company background, team and how we work - [Industries](https://privatedevops.com/industries): Sector-specific infrastructure (e-commerce, SaaS, and more) - [Packages](https://privatedevops.com/packages): Productized engagement tiers - [Free DevOps Tools](https://privatedevops.com/tools): Website and server speed tests, DNS and mail checks, assessments - [Pricing Calculator](https://privatedevops.com/calculator): Estimate infrastructure and engagement cost - [Careers](https://privatedevops.com/careers): Open roles - [Full content map](https://privatedevops.com/llms-full.txt): Expanded version of this file with full service descriptions and FAQs inline