Explore Problems
Showing 190 of 9,941 problems · matching your filters
AI Agents Make Opaque Decisions With No Decision-Level Observability
As AI agents enter production, developers lack tools to trace why an agent made a specific decision rather than just what it did. Traditional APM tools track metrics and logs but not reasoning chains, creating a debugging blindspot. Decision-aware observability is an emerging critical need for reliable agentic systems.
Banks Holding Consumers Liable for Fraudulent Check Fraud in Marketplace Transactions
Banks allow consumers to withdraw funds from deposited checks before they clear, then hold consumers fully liable when checks prove fraudulent. This practice is particularly damaging in peer-to-peer selling contexts where fraudulent payment methods are common. The bank policy of enabling early access while shifting all fraud risk to consumers creates a predictable harm pattern.
AI systems in production lose interpretability as they scale
Engineering teams shipping AI in production report a failure category where standard metrics stay green while the system loses coherence or drifts in non-reproducible ways. The root cause is structural: verification built on the same model that generates creates blind spots that existing observability tooling cannot detect.
Long-Running AI Agent Sessions Require Fragile Shell Multiplexer Workarounds
Developers running long-lived Claude Code or AI agent sessions over SSH must use tmux or screen multiplexers that introduce subtle shell behavior changes and lack standardized safety controls. There is no clean, first-class approach for running multiple parallel isolated agent sessions — a gap that becomes critical as agentic workflows shift toward longer, more autonomous task execution.
No Standard Protocol for AI Agents to Communicate Across Machines
Developers running AI agents on multiple computers or cloud instances have no clean way to route messages between agent instances without custom infrastructure. Existing messaging tools are not designed for agent capability-based discovery. An OSS solution (Viche) emerged using the Erlang actor model to address this gap.
No Standard Protocol for AI Agents to Discover and Compare Real-World Services
AI agents can read web content and call tools but lack a structured way to discover what services a business offers, compare alternatives by SLA and pricing, and place orders autonomously. Existing standards like llms.txt address content readability but not service capability enumeration or procurement workflows. As agents increasingly act as procurement tools, the absence of a machine-readable service manifest format creates a significant integration barrier.
Food Recognition APIs Too Expensive and Inaccurate for Independent Developers
Developers building nutrition or food tracking applications find available food recognition APIs either prohibitively expensive for side projects, unreliable in accuracy, or so poorly documented they are unusable. This forces developers to abandon features or build their own pipelines from scratch. The gap leaves a large class of health and wellness apps unable to add viable food logging.
Mortgage Servicer Double-Charges Property Taxes in Escrow Using Inflated Overlay
LoanCare extracts double the actual county-assessed property tax through escrow by applying a fraudulent administrative neighborhood overlay. The homeowner's county-assessed tax is $3,400 but the servicer charges $6,900 annually, pocketing the difference with no disclosure or justification.
GPU Infrastructure Setup for Robot Physics Simulation is Painful and Repetitive
Robotics engineers setting up GPU-based simulation environments (Isaac Sim, Gazebo, MuJoCo) face significant infrastructure overhead each time they start a new project or join a new team. The process of provisioning, configuring, and tearing down cloud GPU instances for headless simulation runs lacks any CI/CD equivalent, forcing teams to solve the same infra problems repeatedly. The pain is acute enough that teams starting fresh dread the ramp-up, even if they have solved it before.
Brands Have No Visibility Into How AI Engines Mention or Cite Them
As AI-powered search engines (ChatGPT, Perplexity, Gemini) increasingly answer queries instead of directing traffic to websites, brands lose visibility into whether and how they are referenced. There is no established tooling for monitoring brand citations across AI outputs, detecting content gaps, or influencing AI-driven recommendations.
Home insurers cover cosmetic repairs but deny root-cause fixes, then cancel policies
When water damage occurs, insurers pay for interior remediation only — refusing to waterproof the foundation that caused the leak — leaving homeowners with a temporary fix and a recurring problem. The policy language creates a structural gap between what is covered and what constitutes a permanent repair. Insurers compound the harm by cancelling coverage when homeowners document the remediation work that was done.
Salesforce CRM overwhelming feature density drives user abandonment
Salesforce users consistently report feeling overwhelmed by the sheer number of functions, tabs, and options presented without clear hierarchy or guidance. The complexity gap between what most sales teams need and what the platform exposes creates adoption friction. This drives mid-market teams toward lighter CRM alternatives despite Salesforce's feature depth.
European Teams Are Abandoning US SaaS Over Data Privacy and Pricing Risk
GDPR enforcement, the Cloud Act, Schrems II fallout, and volatile USD pricing are pushing European organizations to systematically audit and replace US-based SaaS tools with EU-hosted alternatives. The EU SaaS ecosystem has matured enough to cover most categories including project management, analytics, support, and email. This structural shift creates sustained demand for compliant EU-based alternatives across the entire software stack.
Human Code Review Can't Keep Pace With AI-Generated PR Volume
Engineering teams using AI coding agents now generate far larger, more frequent pull requests than humans can meaningfully review. Teams increasingly lean on automated or AI-assisted review layers to keep production velocity from stalling, raising doubts about how much human oversight remains realistic.
Slack Channel Noise Buries Important Messages as Teams Scale
As team size and channel count grow in Slack, high message volume causes critical communications to get buried under general conversation. Notification overload adds to the problem, and search lacks the contextual ranking needed to surface relevant older messages reliably. Teams have no effective built-in mechanism to separate signal from noise.
Small business owners cannot execute consistent marketing without significant time investment
Small business owners lack the time and marketing expertise to maintain consistent, effective marketing activities. Existing tools require significant learning curves or ongoing manual effort that owners cannot sustain alongside running their business. There is strong demand for solutions that deliver marketing outcomes without requiring owners to become marketers themselves.
No credible open-source bot for automating data-broker removal requests
Paid services exist for opting consumers out of data brokers but feel overpriced or scammy. The repetitive request flow looks well suited to AI automation, yet there is no widely-adopted open-source alternative.
AI Coding Agents Lose Context on Session Reset and Make Opaque Decisions
AI coding assistants forget all reasoning, design decisions, and open TODOs when a session ends, forcing developers to re-explain context from scratch. Compounding this, AI-generated code changes are opaque — it is unclear which prompt or reasoning step caused any given edit. These two gaps block AI agents from functioning as reliable, auditable collaborators in real development workflows.
Long-running coding agents lose task state when context windows overflow or sessions end
Coding agents handling multi-phase tasks store all intermediate state in volatile session context. When context overflows or sessions terminate, the agent loses the full decision history, leading to repeated mistakes and failed handoffs across phases. There is no standard mechanism for externalizing agent workflow state to durable structured storage.
AI security evaluation corrupted by using AI to grade AI outputs
Security practitioners evaluating AI systems face a methodological trap: using AI judges to assess AI behavior introduces circular bias and unreliable verdicts. Human review at scale is impractical, and automated benchmarks do not capture adversarial edge cases. This gap leaves AI deployments with false confidence in their security posture.