Wallow Troubleshooting Guide

This guide helps you diagnose and resolve common issues when developing with Wallow. It covers infrastructure, authentication, database, messaging, testing, build and frontend problems.


Table of Contents

  1. Infrastructure Issues
  2. Authentication Issues
  3. Database Issues
  4. Messaging Issues
  5. Test Failures
  6. Build Issues
  7. Frontend Issues
  8. Debugging Tips

1. Infrastructure Issues

Docker Containers Not Starting

Symptom

docker compose up -d
# Containers exit immediately or show "Restarting" status

Diagnosis

# Check container status
cd docker && docker compose ps

# View logs for specific container
docker compose logs postgres
docker compose logs valkey

Common Causes and Solutions

Port already in use:

Error: bind: address already in use
# Find process using the port (e.g., 5432)
lsof -i :5432

# Kill the process or stop the conflicting service
kill -9 <PID>

# Or change the port in docker-compose.yml

Volume permission issues:

# Reset volumes (WARNING: deletes all data)
cd docker && docker compose down -v
docker compose up -d

Out of disk space:

# Check Docker disk usage
docker system df

# Clean up unused resources
docker system prune -a --volumes

Environment file missing:

# Create .env file from example
cp docker/.env.example docker/.env

PostgreSQL Connection Failures

Symptom

Npgsql.NpgsqlException: Failed to connect to 127.0.0.1:5432
  ---> System.Net.Sockets.SocketException: Connection refused

Diagnosis

# Check if PostgreSQL container is running
docker compose ps postgres

# Test connectivity
docker exec wallow-postgres pg_isready -U wallow

# Check logs
docker compose logs postgres

Solutions

Container not running:

cd docker && docker compose up -d postgres

Wrong connection string: Check appsettings.Development.json or environment variables:

{
  "ConnectionStrings": {
    "DefaultConnection": "Host=localhost;Port=5432;Database=wallow;Username=wallow;Password=wallow"
  }
}

Database not initialized:

# Recreate with init scripts
cd docker && docker compose down -v
docker compose up -d postgres

PostgreSQL not accepting connections:

FATAL: no pg_hba.conf entry for host

Check that the database user has proper permissions. Module schemas are not created by the API on startup — Wallow.MigrationService applies them, and the API only migrates inline in the Testing environment, where Testcontainers hands it an empty database. Under Aspire (pnpm backend), in the e2e stack and in production, the migration service runs to completion first and everything else waits on it.

Valkey/Redis Connection Problems

Symptom

StackExchange.Redis.RedisConnectionException: It was not possible to connect to the redis server(s)

Diagnosis

# Check Valkey container
docker compose ps valkey

# Test connectivity
docker exec wallow-valkey valkey-cli ping
# Should return: PONG

# Check logs
docker compose logs valkey

Solutions

Container not running:

cd docker && docker compose up -d valkey

Wrong connection string: the dev stack's Valkey requires a password, so a bare localhost:6379 authenticates as nobody and the connection is refused. appsettings.Development.json ships the working value:

{
  "ConnectionStrings": {
    "Redis": "localhost:6379,password=WallowValkey123!,abortConnect=false"
  }
}

Memory limit exceeded:

# Check memory usage
docker exec wallow-valkey valkey-cli info memory

# Clear cache if needed
docker exec wallow-valkey valkey-cli FLUSHALL

TLS/SSL configuration issues: For production with TLS:

"Redis": "localhost:6379,ssl=true,abortConnect=false"

API Returns 503 for Everything (First-Run Setup Mode)

Symptom

On a fresh deployment, nearly every request answers 503 Service Unavailable as application/problem+json whose code is Setup.Required — even though every container is healthy. The title and detail are the generic 5xx wording; only the code distinguishes the setup lock from a real outage. Only /v1/identity/setup, /health, /.well-known, /connect, /openapi, and /scalar respond normally.

Cause

This is not an outage. The production seed deliberately creates no administrator, so SetupMiddleware (api/src/Wallow.Api/Middleware/SetupMiddleware.cs) locks the API until one exists.

Solution

Check the status probe, then create the bootstrap admin — via the auth app's setup page or POST /v1/identity/setup/admin:

curl https://your-domain/api/v1/identity/setup/status
# → {"setupRequired": true}

The 503s stop as soon as the admin exists. See the deployment guide's Setup mode section for the full request.


2. Authentication Issues

JWT Validation Failures

Symptom

Microsoft.IdentityModel.Tokens.SecurityTokenSignatureKeyNotFoundException: IDX10500: Signature validation failed

or

401 Unauthorized
WWW-Authenticate: Bearer error="invalid_token"

Diagnosis

# Check that the API is running and healthy
curl http://localhost:5001/health/ready

Solutions

Wrong authentication configuration: Check appsettings.json for correct OpenIddict settings.

Clock skew between server and client:

IDX10222: Lifetime validation failed. The token is expired.

Ensure system clocks are synchronized. JWT has a 5-minute tolerance by default.

Token Expiration Problems

Symptom

Token expired at [timestamp]

Solutions

All tokens come from the OpenIddict token endpoint, POST /connect/token. It takes application/x-www-form-urlencoded parameters, not JSON. There is no email/password token endpoint: the API supports the authorization code (with PKCE), refresh token, and client credentials grants only, so a browser user re-authenticates by going back through /connect/authorize.

Get a fresh token (service account / client credentials):

curl -s -X POST http://localhost:5001/connect/token \
  -d "grant_type=client_credentials" \
  -d "client_id=<your-client-id>" \
  -d "client_secret=<your-client-secret>" \
  -d "scope=inquiries.read inquiries.write"

Use a refresh token:

curl -s -X POST http://localhost:5001/connect/token \
  -d "grant_type=refresh_token" \
  -d "refresh_token=YOUR_REFRESH_TOKEN" \
  -d "client_id=<your-client-id>" \
  -d "client_secret=<your-client-secret>"

Refresh tokens are rolling: issuing a new one revokes the old one, so a retry with an already-used refresh token fails once the reuse leeway (OpenIddict:RefreshTokenReuseLeewaySeconds, default 30) has passed — a later replay revokes the whole token family. Sliding expiration is disabled, meaning the refresh window does not extend on use. Access tokens default to 15 minutes (OpenIddict:AccessTokenLifetimeMinutes). Refresh-token lifetime is per client: the client's own refreshTokenLifetime if set, else 7 days for a seeded first-party client, 1 day for any other application, with OpenIddict:RefreshTokenLifetimeDays (default 7) as the fallback for clients that carry no per-client value — see the Configuration guide.

Missing Claims/Permissions

Symptom

403 Forbidden
{
  "type": "about:blank",
  "title": "Forbidden",
  "status": 403,
  "detail": "The authenticated identity lacks the permission this resource requires.",
  "code": "Auth.Forbidden",
  "traceId": "00-abc123def456...-01"
}

Diagnosis

Decode your JWT token at https://jwt.io and check:

  • role claim - Should contain role names
  • organization claim - Should contain tenant ID

Solutions

User missing role: Assign the required roles via the Identity module's user management API.

Permission not mapped to role: Check PermissionExpansionMiddleware and role-to-permission mappings in: api/src/Modules/Identity/Wallow.Identity.Infrastructure/Authorization/PermissionExpansionMiddleware.cs

Organization claim missing: Ensure user belongs to an organization via the Identity module's organization management API.

Tenant Resolution Failures

Symptom

There is no dedicated exception type for this — the observable is ITenantContext.IsResolved returning false in a handler, and whatever the handler does next when it has no tenant (usually a null-reference or an empty result set where rows were expected).

Diagnosis

Check ITenantContext.IsResolved in your handler before reading TenantId.

Solutions

Missing organization claim: The TenantResolutionMiddleware reads tenant from JWT organization claim or X-Tenant-Id header.

Ensure your JWT contains:

{
  "organization": "00000000-0000-0000-0000-000000000001"
}

Raw SQL without a tenant filter: EF Core's tenant query filters do not apply to FromSql/ExecuteSql, so raw SQL must filter by tenant itself:

await dbContext.Announcements
    .FromSql($"SELECT * FROM announcements.announcements WHERE tenant_id = {_tenantContext.TenantId.Value}")
    .ToListAsync(cancellationToken);

Test environment: In tests, WallowApiFactory registers a fixed tenant context. If you need a different tenant, use the test headers:

client.DefaultRequestHeaders.Add("X-Tenant-Id", "your-tenant-guid");

3. Database Issues

Migration Conflicts

Symptom

Microsoft.EntityFrameworkCore.DbUpdateException: An error occurred while saving the entity changes
  ---> Npgsql.PostgresException: 42P01: relation "identity.users" does not exist

Diagnosis

# Check migration status
dotnet ef migrations list \
  --project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
  --startup-project api/src/Wallow.Api \
  --context IdentityDbContext

Solutions

Apply pending migrations:

dotnet ef database update \
  --project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
  --startup-project api/src/Wallow.Api \
  --context IdentityDbContext

Migration history mismatch:

# Reset database (WARNING: deletes all data)
cd docker && docker compose down -v
docker compose up -d postgres

# Reapply every module's migrations. The API does NOT do this on startup —
# starting it against the empty database just fails differently.
dotnet run --project api/src/Wallow.MigrationService

Conflicting migration:

The migration '20260215_AddNewField' has already been applied to the database
# Remove the conflicting migration
dotnet ef migrations remove \
  --project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
  --startup-project api/src/Wallow.Api \
  --context IdentityDbContext

EF Core Tracking Issues

Symptom

System.InvalidOperationException: The instance of entity type 'Invoice' cannot be tracked because another instance with the same key value is already being tracked

Solutions

Use AsNoTracking for read-only queries:

var invoices = await _context.Invoices
    .AsNoTracking()
    .Where(i => i.Status == InvoiceStatus.Pending)
    .ToListAsync();

Detach existing entity:

var existingEntry = _context.Entry(existingInvoice);
existingEntry.State = EntityState.Detached;

Use new DbContext scope:

using var scope = _serviceProvider.CreateScope();
var context = scope.ServiceProvider.GetRequiredService<IdentityDbContext>();
// Now you have a fresh tracking context

Connection Pool Exhaustion

Symptom

Npgsql.NpgsqlException: The connection pool has been exhausted
  ---> System.InvalidOperationException: Timeout expired. The timeout period elapsed prior to obtaining a connection from the pool.

Diagnosis

-- Check active connections
SELECT count(*) FROM pg_stat_activity WHERE datname = 'wallow';

-- See connection details
SELECT pid, usename, application_name, state, query_start
FROM pg_stat_activity
WHERE datname = 'wallow'
ORDER BY query_start DESC;

Solutions

Increase pool size:

{
  "ConnectionStrings": {
    "DefaultConnection": "Host=localhost;...;Maximum Pool Size=100;Connection Idle Lifetime=300"
  }
}

Dispose connections properly:

// Use 'using' or 'await using' for DbContext
await using var context = await _contextFactory.CreateDbContextAsync();

Close long-running connections:

-- Terminate idle connections older than 10 minutes
SELECT pg_terminate_backend(pid)
FROM pg_stat_activity
WHERE datname = 'wallow'
  AND state = 'idle'
  AND query_start < now() - interval '10 minutes';

4. Messaging Issues

Wallow uses Wolverine with in-memory messaging. There is no external message broker.

Messages Not Being Delivered

Symptom

  • Events published but handlers never execute
  • No errors in logs

Solutions

Handler not discovered: Wolverine discovers handlers only in the assemblies each enabled module declares through IWallowModule.HandlerAssemblies. Ensure:

  1. Handler class is public static
  2. Method is named Handle or HandleAsync
  3. The handler's assembly is listed in its module's HandlerAssemblies (every module declares both its .Application and its .Infrastructure assembly, so a handler in either is already covered)
  4. The owning module is enabled — a module switched off in FeatureManagement:Modules.* contributes no handler assemblies at all, so its messages go unhandled
// Correct handler pattern
public static class MyEventHandler
{
    public static async Task HandleAsync(MyEvent @event, IMyService service, CancellationToken ct)
    {
        // Handle event
    }
}

Check Wolverine envelope tables for errored messages:

SELECT * FROM wolverine.wolverine_incoming_envelopes WHERE status = 'error';

Handler Not Being Discovered

Symptom

Wolverine.Runtime.UnknownMessageTypeException: Unknown message type 'MyEvent'

Diagnosis

Check Wolverine's discovered handlers at startup in logs.

Solutions

Ensure handler follows conventions:

// Must be public static class
public static class MyEventHandler
{
    // Method must be Handle or HandleAsync
    // First parameter must be the message type
    public static async Task HandleAsync(MyEvent @event, ILogger<MyEvent> logger)
    {
        // ...
    }
}

Assembly not included in discovery: Program.cs does not scan for assemblies — it hands Wolverine exactly what the enabled modules declare:

Assembly[] handlerAssemblies =
[
    typeof(Wallow.Api.WallowModules).Assembly,   // the host
    typeof(IWallowModule).Assembly,              // Wallow.Shared.Infrastructure — no module owns it
    .. enabledModules.SelectMany(module => module.HandlerAssemblies),
];

So the fix is in the owning module, not in Program.cs. Add the assembly to that module's HandlerAssemblies in api/src/Modules/{Module}/Wallow.{Module}.Infrastructure/Modules/{Module}Module.cs:

public IReadOnlyList<Assembly> HandlerAssemblies =>
[
    typeof(CreateThingHandler).Assembly,   // .Application
    typeof({Module}Module).Assembly,       // .Infrastructure
    typeof(MyExtraHandler).Assembly,       // the assembly that was missing
];

A brand-new module also needs its entry in WallowModuleRegistry.All (api/src/Wallow.Modules.Registry/WallowModuleRegistry.cs) — nothing else discovers it.

Outbox Not Processing

Symptom

Messages stuck in wolverine.wolverine_outgoing_envelopes table.

Diagnosis

-- Check outbox status
SELECT status, count(*)
FROM wolverine.wolverine_outgoing_envelopes
GROUP BY status;

-- View stuck messages
SELECT * FROM wolverine.wolverine_outgoing_envelopes
WHERE status = 'scheduled'
ORDER BY scheduled_time;

Solutions

Wolverine agent not running: The Wolverine durability agent processes the outbox. Ensure it's enabled:

opts.PersistMessagesWithPostgresql(connectionString, "wolverine");

Database transaction not committed: Messages are only sent when the transaction commits:

// Ensure SaveChanges is called
await _context.SaveChangesAsync();
// Outbox messages are now ready for sending

Agent polling interval: By default, Wolverine polls every 5 seconds. For debugging:

opts.Durability.PollingInterval = TimeSpan.FromSeconds(1);

5. Test Failures

Testcontainers Not Starting

Symptom

Docker.DotNet.DockerApiException: Docker API responded with status code=InternalServerError

or

Testcontainers.Containers.ContainerStartException: The container did not start in time

Diagnosis

# Verify Docker is running
docker info

# Check Docker resources
docker system info | grep -E "CPUs|Memory"

Solutions

Docker not running: Start Docker Desktop or Docker daemon.

Insufficient resources: In Docker Desktop settings, increase:

  • Memory: At least 4GB
  • CPUs: At least 2

Port conflicts:

# Find conflicting ports
lsof -i :5432
lsof -i :6379

Container image not found:

# Pull images manually
docker pull postgres:18-alpine
docker pull valkey/valkey:8-alpine

Timeout too short:

// Increase wait time in test fixture
private readonly PostgreSqlContainer _postgres = new PostgreSqlBuilder()
    .WithWaitStrategy(Wait.ForUnixContainer()
        .UntilPortIsAvailable(5432)
        .WithTimeout(TimeSpan.FromMinutes(2)))
    .Build();

Parallel Test Conflicts

Symptom

Tests pass individually but fail when run together:

System.InvalidOperationException: Database is in use by another process

Solutions

Use test collection to run sequentially:

[Collection("Database")]
public class InvoiceTests : IClassFixture<WallowApiFactory>
{
    // Tests in same collection run sequentially
}

Isolate test data:

// Use unique identifiers per test
var invoiceId = Guid.NewGuid();
var tenantId = Guid.NewGuid();

Disable parallel execution: In xunit.runner.json:

{
  "parallelizeTestCollections": false
}

SignalR Test Issues

Symptom

System.IO.IOException: The server returned status code '401' when status code '101' was expected

Solutions

Include auth token in connection:

var connection = new HubConnectionBuilder()
    .WithUrl($"{_factory.Server.BaseAddress}hubs/realtime", options =>
    {
        options.AccessTokenProvider = () => Task.FromResult<string?>(_token);
        options.HttpMessageHandlerFactory = _ => _factory.Server.CreateHandler();
    })
    .Build();

Use TestAuthHandler headers:

// The TestAuthHandler reads auth from query parameters too
var url = $"{baseUrl}hubs/realtime?access_token={token}";

Wait for connection:

await connection.StartAsync();
// Give SignalR time to establish connection
await Task.Delay(500);

Integration Test Authentication

Symptom

Tests return 401 even with TestAuthHandler.

Solutions

Ensure test scheme is used: WallowApiFactory configures a "Test" authentication scheme:

services.AddAuthentication("Test")
    .AddScheme<AuthenticationSchemeOptions, TestAuthHandler>("Test", options => { });

Set required headers:

client.DefaultRequestHeaders.Add("X-Test-User-Id", Guid.NewGuid().ToString());
client.DefaultRequestHeaders.Add("X-Test-Roles", "admin");

Skip auth for specific tests:

client.DefaultRequestHeaders.Add("X-Test-Auth-Skip", "true");

6. Build Issues

Package Restore Failures

Symptom

error NU1101: Unable to find package Wallow.Storage.Domain

Solutions

Restore from solution root:

cd /path/to/Wallow
dotnet restore

Clear NuGet cache:

dotnet nuget locals all --clear
dotnet restore

Check package sources:

dotnet nuget list source
# Ensure nuget.org is present

Verify network connectivity:

curl https://api.nuget.org/v3/index.json

Project Reference Problems

Symptom

error CS0246: The type or namespace name 'AnnouncementDto' could not be found

Solutions

Check project references:

# View project references
dotnet list api/src/Modules/Announcements/Wallow.Announcements.Api/Wallow.Announcements.Api.csproj reference

Add missing reference:

dotnet add api/src/Modules/Announcements/Wallow.Announcements.Api/Wallow.Announcements.Api.csproj \
  reference api/src/Modules/Announcements/Wallow.Announcements.Application/Wallow.Announcements.Application.csproj

Clean and rebuild:

dotnet clean
dotnet build

Assembly Conflicts

Symptom

System.IO.FileLoadException: Could not load file or assembly 'Newtonsoft.Json, Version=13.0.0.0'

Solutions

Check for version conflicts: All package versions are centrally managed in api/Directory.Packages.props:

grep -r "Newtonsoft.Json" api/Directory.Packages.props

Enable binding redirects: In the project file:

<PropertyGroup>
  <AutoGenerateBindingRedirects>true</AutoGenerateBindingRedirects>
</PropertyGroup>

Clear build artifacts:

dotnet clean
rm -rf */bin */obj
dotnet build

7. Frontend Issues

Wallow is a polyglot monorepo: the sections above cover the .NET half, this one covers the pnpm workspace under apps/ and packages/.

pnpm lint passes but CI fails on a test file

pnpm lint lints source only — it excludes **/*.test.* and **/*.stories.tsx. The excluded files are linted by the second pass, pnpm lint:tests, which additionally enables oxlint's vitest plugin. Together the two cover every file exactly once, and pnpm check runs both. Running only the first one and concluding you are clean is the usual cause.

pnpm lint          # source
pnpm lint:tests    # test + story files, with the vitest plugin

scripts/lint-tests.sh fails loudly if it enumerates zero files, because oxlint does not expand globs in path arguments and a silent zero-file pass looks exactly like success.

A change to package source does not take effect

build, typecheck and test run through turbo with content-addressed caching in .turbo/. If a task replays a stale result, the input you changed is not in that task's hash. Force a run to confirm:

pnpm build --force            # bypass the cache for one run
rm -rf .turbo                 # or drop the local cache entirely

Caching is local unless TURBO_API/TURBO_TEAM/TURBO_TOKEN are exported — with them set, turbo also reads and writes the shared remote cache, so a stale entry can come from another machine (see the developer guide's Turbo remote cache section; TURBO_REMOTE_CACHE_READ_ONLY=true or unsetting TURBO_TOKEN takes the remote out of the picture). Note that lint, format, manifest, dependency and export checks are not turbo tasks; they are root scripts and always run.

Vitest fails to launch a browser

DOM specs run in real headless Chromium through Vitest browser mode with the Playwright provider, never jsdom or happy-dom. A missing browser binary shows up as a launch failure before any test runs:

pnpm exec playwright install chromium

If a spec that touches the DOM instead runs in Node — document is not defined, or assertBrowserModeSmoke failing — the file is in the wrong vitest project. Check the consumer's vite.config.ts project inventory; see packages/testing/CLAUDE.md for the split.

E2E suites fail against a backend that is not there

The Playwright suites in apps/wallow-auth/e2e/, apps/wallow-web/e2e/ and apps/wallow-web/e2e-cross-app/ are backend-dependent. Run them through the one-command runner, which brings the containerised stack up and tears it down again:

./scripts/e2e.sh

Running pnpm exec playwright test directly against a stack you have not started fails at the first navigation.


8. Debugging Tips

Enabling Detailed Logging

Serilog configuration:

{
  "Serilog": {
    "MinimumLevel": {
      "Default": "Debug",
      "Override": {
        "Microsoft": "Information",
        "Microsoft.EntityFrameworkCore": "Warning",
        "Wolverine": "Debug"
      }
    }
  }
}

EF Core query logging:

{
  "Logging": {
    "LogLevel": {
      "Microsoft.EntityFrameworkCore.Database.Command": "Information"
    }
  }
}

Wolverine message logging: Already configured in Program.cs:

opts.ConfigureMessageLogging(); // Logs message execution

Checking Wolverine Envelope Tables

View pending outbox messages:

SELECT
    id,
    message_type,
    destination,
    status,
    scheduled_time,
    attempts
FROM wolverine.wolverine_outgoing_envelopes
ORDER BY scheduled_time DESC
LIMIT 50;

View incoming (inbox) messages:

SELECT
    id,
    message_type,
    status,
    received_at
FROM wolverine.wolverine_incoming_envelopes
ORDER BY received_at DESC
LIMIT 50;

Clear stuck messages:

-- WARNING: This may lose messages
DELETE FROM wolverine.wolverine_outgoing_envelopes WHERE status = 'error';

Quick Diagnostic Commands

Check all services:

cd docker && docker compose ps
docker compose logs --tail=50

Test database connectivity:

docker exec wallow-postgres psql -U wallow -d wallow -c "SELECT 1"

Test Redis/Valkey connectivity:

docker exec wallow-valkey valkey-cli ping

View application logs:

# If running via dotnet run
# Logs output to console

# If running in Docker
docker logs wallow-api --tail=100 -f

Reset everything:

cd docker && docker compose down -v && docker compose up -d

Quick Reference: Error Code Mapping

HTTP Status Exception Type Meaning
400 ValidationException, ArgumentException Invalid request data
401 UnauthorizedAccessException Authentication failed
403 ForbiddenAccessException Missing permission
404 EntityNotFoundException Resource not found
422 BusinessRuleException Business rule violation
500 Unhandled exception Server error

Getting Help

If you cannot resolve an issue:

  1. Check existing documentation:

    • docs/getting-started/developer-guide.md - Development setup
    • docs/operations/deployment.md - Production deployment
  2. Search the codebase:

    grep -r "error message" api/src/
    
  3. Check recent commits:

    git log --oneline -20
    
  4. Run architecture tests:

    ./scripts/run-tests.sh arch