<?xml version="1.0" encoding="utf-8" standalone="yes"?><?xml-stylesheet type="text/xsl" href="https://bytejournal.org/feed.xsl"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Infrastructure — Byte Journal</title><link>https://bytejournal.org/topics/infrastructure/</link><description>Technology news and writing on AI, infrastructure and security.</description><language>en</language><generator>Hugo 0.167.0</generator><ttl>60</ttl><atom:link href="https://bytejournal.org/topics/infrastructure/index.xml" rel="self" type="application/rss+xml"/><lastBuildDate>Fri, 02 Oct 2026 12:15:01 -0400</lastBuildDate><item><title>Cloudflare launches Workers KV Instant in private beta</title><link>https://bytejournal.org/news/2026-10-01-cloudflare-launches-workers-kv-instant-in-private-beta/</link><guid isPermaLink="true">https://bytejournal.org/news/2026-10-01-cloudflare-launches-workers-kv-instant-in-private-beta/</guid><pubDate>Fri, 02 Oct 2026 12:15:01 -0400</pubDate><dc:creator>Arxchitect</dc:creator><description>Cloudflare launched Workers KV Instant, a new mode for its edge key-value store backed by its internal Quicksilver engine, into private beta on Oct 1, claiming sub-2ms p99 reads and reads priced 60% lower than classic KV.</description><content:encoded><![CDATA[<p>Cloudflare pulled a key-value store out of its internal toolkit, opening Workers KV Instant, a new mode for its Workers KV service, to a private beta. The company&rsquo;s announcement frames the offering around two claims: data written can be pushed across its network almost immediately, and the &ldquo;cold read&rdquo; penalty that has long hit developers fetching a key from an edge location for the first time is removed.</p>
<p>The headline numbers are Cloudflare&rsquo;s own. Read latency in Instant mode resolves in under two milliseconds at the 99th percentile, which the company says is more than one hundred times faster than classic mode, with 95th-percentile reads measured in microseconds. Spanish tech outlet El Ecosistema Startup, which covered the same launch the day it broke, pulled the underlying comparison table from Cloudflare&rsquo;s post: p99 reads at 1.62ms for Instant against 287ms for classic (160ms cached), median write replication at 107ms, p99 write replication at 256ms versus 4.38 seconds for classic, and roughly 99% of writes reaching all edge locations in about 250ms. Cloudflare operates more than 300 edge locations.</p>
<p>The noteworthy part is the engine underneath. Cloudflare says KV Instant runs on Quicksilver v2, the globally distributed key-value system it built internally and has relied on since around 2020, noting that nearly every request to its network performs at least one key lookup. Independent technical outlet InfoQ reported in 2025 on how Cloudflare migrated that store from an architecture that kept the full dataset on every server to a tiered caching model, adding local caches, data-center sharded caches, and full replicas on dedicated storage nodes backed by RocksDB, a change driven partly by unsustainable growth in the dataset.</p>
<p>For developers, the pitch is that existing Workers KV code carries over. The familiar get(), put(), list() and delete() API remains, with Cloudflare noting just three operational differences: you must declare Instant mode when creating a namespace, metadata is unsupported (getWithMetadata always returns null), and list operations return all matching keys without pagination. The company also imposes a one-write-per-namespace-per-second ceiling, which it likens to classic KV&rsquo;s per-key write limit and frames as helping reason about update ordering.</p>
<p>Pricing is a headline differentiator. Cloudflare prices reads (class B operations) at $0.20 per million, sixty percent below classic KV&rsquo;s $0.50, with class A writes billed per operation at $0.10. The trade-off is storage: Instant is billed at $100 per megabyte per month versus $0.50 per gigabyte for classic. Constraints are tight, with keys capped at 300 bytes, values sized so a namespace stays under one megabyte total, and up to 10,000 key-value pairs per namespace.</p>
<p>Caveats matter here. This is a private beta, so none of the performance figures are independently measured, and most operational detail is unverified. Independent reporting on this specific launch was thin: beyond Cloudflare&rsquo;s own post and the Spanish outlet&rsquo;s coverage, much of what circulates online is a repost of the announcement. InfoQ&rsquo;s piece predates the launch and examines the underlying store rather than KV Instant itself. What is clear is that Cloudflare is exposing the store behind its own infrastructure to customers, at pricing that favors very frequent, very fast reads over cheap bulk storage.</p>
]]></content:encoded><category>infrastructure</category><category>Infrastructure</category><source url="https://blog.cloudflare.com/workers-kv-instant/">Cloudflare</source></item><item><title>Cloudflare Basin data platform reaches general availability</title><link>https://bytejournal.org/news/2026-10-01-cloudflare-basin-data-platform-reaches-general-availability/</link><guid isPermaLink="true">https://bytejournal.org/news/2026-10-01-cloudflare-basin-data-platform-reaches-general-availability/</guid><pubDate>Fri, 02 Oct 2026 11:15:01 -0400</pubDate><dc:creator>Arxchitect</dc:creator><description>Cloudflare's Basin data analytics platform is generally available, rebranded from the Cloudflare Data Platform and built on Apache Iceberg and R2 object storage with free egress for data leaving the platform.</description><content:encoded><![CDATA[<p>Cloudflare formally released its data analytics suite, promoting it to general availability and giving the whole platform a new name: Basin. As the company explained in its announcement, Basin is a serverless stack for collecting, storing, and querying analytical data, and it absorbs the three products that made up the old beta program, which had shipped under the name &ldquo;Cloudflare Data Platform&rdquo; during Birthday Week 2025. Basin Pipelines, formerly Cloudflare Pipelines, ingests and SQL-transforms events before writing them to object storage. Basin Catalog, formerly R2 Data Catalog, handles the metadata layer for tables. Basin SQL, formerly R2 SQL, is the serverless query engine.</p>
<p>For people who run servers, the pitch is essentially an observational data layer that doesn&rsquo;t ask you to stand up and babysit compute. In press materials, Cloudflare CTO Dane Knecht framed Basin as an alternative to assembling your own hardware or clusters to observe data, per The Register&rsquo;s coverage. Telecompaper, which also carried the launch, described the platform the same way: something that pulls in data from applications, infrastructure devices, and other Cloudflare services and turns it into production-grade workflows.</p>
<p>The two technical pillars Basin leans on are worth separating. The first is openness. Everything sits on Apache Iceberg, the open table format, with storage in Cloudflare&rsquo;s R2 object storage. Cloudflare says you can read and write your data with any Iceberg-compatible engine — it names DuckDB, PyIceberg, Snowflake, and Apache Spark — which keeps storage decoupled from compute and avoids dumping you into a single vendor&rsquo;s query tooling. The Register points out that rival hyperscaler offerings such as Amazon S3, Google Cloud Storage and Azure Blob Storage charge egress fees, which is where the second pillar comes in.</p>
<p>The second pillar, and the trickier one, is the zero-egress promise. Cloudflare claims that querying or moving data out of R2, including across regions and clouds, carries no data-transfer charge. The Register&rsquo;s coverage is careful to crawl Cloudflare&rsquo;s own fine print: direct egress from R2 via the Workers API, S3 API, or the r2.dev domains is free, but if you connect other metered services to a bucket, those services can bill you on their own, and using Basin Pipelines to land streaming data into Iceberg or Parquet in R2 incurs delivery charges. So &ldquo;no egress fees&rdquo; is real for the storage layer itself, with caveats once other paid services get involved.</p>
<p>Billing, per Cloudflare, is usage-based: you&rsquo;re charged only when Basin ingests, processes, or queries data, with no hourly rates or separate infrastructure costs. The company says rendering data queryable is a matter of seconds, and that Basin SQL scales across its global network by splitting work across Workers. Basin Catalog handles routine table maintenance — compaction, snapshot retention and expiry, manifest clustering — so you don&rsquo;t run a separate Spark job to keep tables healthy.</p>
<p>Cloudflare says early adopters already ran on the platform in beta, including its own billing and infrastructure teams. It cites customers replacing an AWS S3 and Athena setup with Basin and using the zero-egress access to distribute data products across regions. Tens of thousands of Pipelines were reportedly created during the beta, with ingestion scaling to 3 GB/s per stream.</p>
<p>Not yet independently confirmed: real-world performance on very large analytical workloads, exact per-gigabyte pricing, and how the egress fine print plays out in mixed multi-cloud setups. Existing Cloudflare Pipelines, R2 Data Catalog, and R2 SQL configurations continue to work under the new names.</p>
]]></content:encoded><category>infrastructure</category><category>Infrastructure</category><source url="https://blog.cloudflare.com/cloudflare-basin/">Cloudflare</source></item><item><title>Cloudflare launches Auto Router in public beta to cut AI model costs</title><link>https://bytejournal.org/news/2026-10-02-cloudflare-launches-auto-router-in-public-beta-to-cut-ai-model-costs/</link><guid isPermaLink="true">https://bytejournal.org/news/2026-10-02-cloudflare-launches-auto-router-in-public-beta-to-cut-ai-model-costs/</guid><pubDate>Fri, 02 Oct 2026 11:01:12 -0400</pubDate><dc:creator>Arxchitect</dc:creator><description>Cloudflare's Auto Router, now in public beta via AI Gateway, routes requests to cheaper models and claims up to 30% cost savings in internal tests.</description><content:encoded><![CDATA[<p>Cloudflare released its Auto Router in public beta through AI Gateway, a service that automatically sends each AI request to a model capable enough for the task. The company says the tool can cut costs by up to 30% compared with using only frontier models, based on its internal use through the OpenCode harness. The release matters because organizations adopting AI face rising token spend, and the router aims to reduce that spend without requiring users to pick models manually.</p>
<p>Users set their model to cloudflare/auto, and the router handles selection. In an internal benchmark of general knowledge work, cloudflare/auto achieved an 86.6% success rate across 252 of 291 trials, at a total cost of $2.10 and $0.0084 per success. Anthropic&rsquo;s Claude Opus 5.5 scored 96.6% (281/291) at $5.91 total and $0.0210 per success, while OpenAI&rsquo;s GPT-6 Sol scored 84.2% (245/291) at $2.64 total and $0.0108 per success. The benchmark used 97 tasks with three samples per model per task, and confidence intervals were estimated from 10,000 task-level bootstrap resamples. Cloudflare says Auto Router came in at 80% the cost of Sol and 35% the cost of Opus.</p>
<p>The router works in two stages. It first builds a pool of models that can serve the request, filtering by format, execution mode, credentials, billing, access policies, spend limits, and provider health. It then sends a compact view of the conversation to a multi-head classification model running on Workers AI and deployed on GPUs across Cloudflare&rsquo;s edge network. That classifier assigns probabilities across 14 task categories and rates the request on complexity, ambiguity, stakes, and dependence on earlier context, each on a one-to-five scale. A scoring matrix combines those signals with model benchmark results to estimate fit, and the router selects the model with the highest utility, defined as expected quality minus an adaptive cost penalty.</p>
<p>Cloudflare says the design makes routing decisions legible, because users can inspect a task&rsquo;s predicted category and complexity. Adding a new model does not require retraining; the company only adds benchmark-derived weights to the scoring matrix. The router also accounts for cache costs in long agentic sessions. Within a turn, it rarely switches models because the cache is hot. Across turns, it applies a switching penalty that grows with the number of tokens already in context, pricing a model with a live cache at its cheaper cache-read rate and every other candidate at the full cost of rewriting the context. Cloudflare notes that most models cannot read another model&rsquo;s reasoning tokens, so a switch that drops reasoning tokens may force the new model to redo that work at output prices. In the future, the company wants the router to prefer staying within the same model family when it switches.</p>
<p>Cloudflare positions AI Gateway as a control plane for organizations deploying AI internally. Because every request from users, agents, and tools flows through it, the gateway can do more than observe and enforce budgets, spend limits, and identity-aware analytics. The company says budgets and rules still rely on individuals to make cost-conscious choices request by request, and the next step is for the gateway to make intelligent decisions on a user&rsquo;s behalf. Cloudflare also plans to release other routers, including cloudflare/auto-best, which uses the same classification and model pool but selects the highest expected quality without applying the cost tradeoff.</p>
<p>The public beta is available now through AI Gateway, but several details remain unproven. Cloudflare&rsquo;s cost savings and benchmark results come from its internal usage and its own general knowledge work benchmark, not from independent testing. The company says Auto Router does best across a wide range of knowledge-work tasks, like those in a large organization spanning technical and non-technical teams, and that its internal results for coding tasks are comparable with frontier models. Whether those results hold for other organizations, workloads, or model pools is not yet established. The release is also only the beginning of what Cloudflare says the Auto Router can learn from its position in the inference path.</p>
]]></content:encoded><category>infrastructure</category><category>Infrastructure</category><source url="https://blog.cloudflare.com/auto-router/">Cloudflare</source></item><item><title>Fastly and Google Cloud launch an autonomous edge security agent for Gemini Enterprise</title><link>https://bytejournal.org/news/2026-10-01-fastly-and-google-cloud-launch-autonomous-edge-defense-agent-in-gemini/</link><guid isPermaLink="true">https://bytejournal.org/news/2026-10-01-fastly-and-google-cloud-launch-autonomous-edge-defense-agent-in-gemini/</guid><pubDate>Thu, 01 Oct 2026 14:03:43 -0400</pubDate><dc:creator>Arxchitect</dc:creator><description>Fastly and Google Cloud bring an AI security agent for edge incidents into Gemini Enterprise, pairing edge telemetry with Gemini reasoning to cut mean time to resolution from hours to seconds.</description><content:encoded><![CDATA[<p>Fastly has teamed up with Google Cloud to put an AI security agent for edge and infrastructure incidents directly inside Gemini Enterprise. The new offering, called the Autonomous Edge Defense Agent (AEDA), is meant to speed up the way teams triage and fix production problems. Fastly&rsquo;s blog says it can cut mean time to resolution (MTTR) from hours down to seconds.</p>
<p>Fastly describes AEDA as resting on a two-context architecture that pairs the company&rsquo;s live request telemetry, global traffic patterns, and edge control plane with Gemini Enterprise&rsquo;s stateful reasoning. It blends structured network telemetry with unstructured threat intelligence, and Google Cloud says it lets security teams investigate edge and infrastructure incidents in plain language inside Gemini Enterprise instead of manually parsing logs.</p>
<p>When a user flags an anomaly, AEDA can run an active root-cause-analysis loop, according to Fastly. The loop first generates hypotheses about plausible attack paths based on the initial edge signals, then executes autonomous background calls to Fastly&rsquo;s global data layer, origin infrastructure, and telemetry stores. It finishes by translating the findings into remediation recommendations such as rate-limiting rules, WAF signature updates, or origin shielding changes, each delivered with supporting evidence, expected impact, and step-by-step guidance on how to apply it.</p>
<p>A central selling point is separating a coordinated attack from a routine application bug. Fastly says AEDA cross-references an organization&rsquo;s telemetry against anonymized, aggregated global traffic to tell whether an anomaly is isolated or part of a broader campaign. Google Cloud frames this the same way, noting the agent pairs an organization&rsquo;s own telemetry with anonymized intelligence drawn from Fastly&rsquo;s global customer base to determine in seconds whether an incident is isolated or part of a larger attack on the sector, returning evidence-backed findings with a recommended fix.</p>
<p>The launch isn&rsquo;t a one-off. Google Cloud said at its Cloud Next event that it is expanding the catalog of partner-built security agents in Gemini Enterprise, and AEDA is one of roughly twenty vendor offerings in that batch, alongside agents from Acalvio, Britive, Check Point, CrowdStrike, Cyera, Fortinet, Palo Alto Networks, Qualys, Splunk, Snyk, and Zscaler, among others. The agents fall into two buckets: security agents teams invoke directly in the Gemini Enterprise interface, and protections aimed at AI and agentic workloads. Google Cloud says the agents are discoverable and deployable through the Google Cloud Marketplace.</p>
<p>A couple of caveats are worth flagging. Fastly positions the investigation loop as user-triggered rather than fully hands-off, so a human still initiates the analysis. Neither company&rsquo;s announcement lists general-availability timing, pricing, or regional rollout details, and the two write-ups are joint product announcements, so the practical scope of what customers can actually deploy today remains to be confirmed.</p>
]]></content:encoded><category>infrastructure</category><category>Infrastructure</category><source url="https://www.fastly.com/blog/fastly-unveils-aeda-autonomous-edge-security-gemini-enterprise/">Fastly</source></item></channel></rss>