Why computers still waste time and energy.
Engineers, operators, researchers, and technical buyers
Valen Systems studies the scheduling, memory, accelerator, and coordination problems that determine how much useful work a machine can do.
The Scheduler Can’t See the Work
Why a general-purpose scheduler sees runnable work but not the request, dependency, cache, or accelerator context behind it.
→An Idle CPU Isn’t Always the Right CPU
Why CPU placement has to account for cache, NUMA distance, capacity, and cooperating work instead of idle time alone.
→Migration Trades Locality for Balance
Why moving work can restore balance while discarding cache and memory locality that the task had already earned.
→Shorter Slices Don’t Create Free Responsiveness
Why responsiveness and throughput pull scheduler decisions in different directions.
→Fair CPU Time Isn’t Business Importance
Why fair-share accounting can’t by itself know which work matters most to a person or a business.
→History Arrives After the Work Changes
Why historical utilization is useful for prediction but can lag behind sudden workload changes.
→Idle Doesn’t Mean Equivalent
Why performance cores, efficiency cores, cache domains, and capacity differences complicate CPU choice.
→Energy Models Are Models
What energy-aware scheduling can predict, and why the prediction depends on hardware models and workload conditions.
→Finding Local Memory Costs Something
Why automatic NUMA balancing may improve locality while imposing scanning, unmapping, fault, and migration work.
→Coordination Gets More Expensive as the Machine Grows
How run queues, scheduling domains, cache topology, NUMA nodes, and cross-core coordination complicate large machines.
→Two Logical CPUs Still Share a Core
Why simultaneous multithreading creates both performance-sharing and security-policy conflicts.
→A vCPU May Hide a Whole Scheduler
Why host scheduling can miss guest runnable state and the cost of delaying a CPU thread that feeds an accelerator.
→One Default Policy Can’t Be Optimal for Every Workload
Why a default scheduler must serve phones, desktops, servers, embedded systems, and large machines without specializing for one workload.
→Local Scheduler Decisions Can Create Global Effects
Why wakeup placement, frequency, NUMA, SMT, cgroups, and device work interact in ways no local rule sees completely.
→sched_ext Makes Experiments Easier, Not Automatically Safe
What Linux sched_ext enables for workload-specific scheduling and why its API and fallback boundaries matter.
→A Warp Takes the Branches Together
Why divergent branches inside a GPU warp serialize paths and reduce effective parallel utilization.
→Occupancy Is a Resource Allocation Problem
Why GPU occupancy is constrained by registers, shared memory, threads, and execution slots rather than one speed setting.
→GPU Memory Wants Neighboring Threads
Why scattered GPU memory accesses create more transactions and reduce effective bandwidth.
→A Busy GPU Can Still Have Idle Compute Units
Why uneven block completion can leave compute resources unused while other blocks continue running.
→Tiny Kernels Pay a Launch Cost
Why batching can improve throughput while making latency and dependency timing less direct.
→Throughput Queues Don’t Behave Like Latency Queues
Why long GPU work can delay latency-sensitive work and why interruption is often more constrained than on a CPU.
→GPU Kernels Share More Than the Compute Units
Why kernels can interfere through memory bandwidth, caches, execution units, and copy paths even when compute capacity remains available.
→GPU Throughput Moves When Power and Temperature Move
Why a scheduler can’t assume a fixed GPU execution rate when clocks and power limits change.
→Independent Kernels Can Overlap. Dependent Kernels Can’t
Why accelerator scheduling becomes a graph problem when kernels share dependencies and data movement.
→The CPU and GPU Can Wait on Each Other
Why poor synchronization can leave either processor idle even when the total system has available capacity.
→More GPUs Add Communication to the Schedule
Why multi-GPU work must account for partitioning, transfer bandwidth, topology, and synchronization.
→Maximum Throughput Isn’t a Real-Time Guarantee
Why predictable GPU latency remains harder than maximizing average throughput.
→DRAM Remembers a Row, Not Your Intent
Why memory controllers care about banks, open rows, and request order when many cores share DRAM.
→The Memory Bus Pays to Change Direction
Why batching reads or writes can reduce bus turnaround cost while increasing waiting time for the other class.
→More Memory Threads Stop Helping at Saturation
Why finite bandwidth turns added concurrency into queueing once the memory system is full.
→Remote Memory Is a Topology Decision
Why a task's CPU placement and memory placement can’t be reasoned about independently on a NUMA machine.
→Memory Has Maintenance Windows
Why DRAM maintenance and refresh behavior complicate the assumption that memory is continuously available at one fixed latency.
→VRAM Contention Is a Shared-Bandwidth Problem
Why compute, rendering, encoding, transfers, and caches can compete for the same device memory system.
→Unified Memory Can Move the Bottleneck
Why page migration between system memory and device memory can trade programming convenience for transfer and fault cost.
→Free Memory Isn’t Always a Usable Block
Why allocation shape, fragmentation, and data-dependent compression affect usable memory capacity and traffic.
→Can Workload Coordination Reduce Memory Contention?
A research question about whether better workload information can reduce memory conflicts and burst pressure.
→Working through a hardware or runtime decision? Explore compute evaluations