Wallow Troubleshooting Guide
This guide helps you diagnose and resolve common issues when developing with Wallow. It covers infrastructure, authentication, database, messaging, testing, build and frontend problems.
Table of Contents
- Infrastructure Issues
- Authentication Issues
- Database Issues
- Messaging Issues
- Test Failures
- Build Issues
- Frontend Issues
- Debugging Tips
1. Infrastructure Issues
Docker Containers Not Starting
Symptom
docker compose up -d
# Containers exit immediately or show "Restarting" status
Diagnosis
# Check container status
cd docker && docker compose ps
# View logs for specific container
docker compose logs postgres
docker compose logs valkey
Common Causes and Solutions
Port already in use:
Error: bind: address already in use
# Find process using the port (e.g., 5432)
lsof -i :5432
# Kill the process or stop the conflicting service
kill -9 <PID>
# Or change the port in docker-compose.yml
Volume permission issues:
# Reset volumes (WARNING: deletes all data)
cd docker && docker compose down -v
docker compose up -d
Out of disk space:
# Check Docker disk usage
docker system df
# Clean up unused resources
docker system prune -a --volumes
Environment file missing:
# Create .env file from example
cp docker/.env.example docker/.env
PostgreSQL Connection Failures
Symptom
Npgsql.NpgsqlException: Failed to connect to 127.0.0.1:5432
---> System.Net.Sockets.SocketException: Connection refused
Diagnosis
# Check if PostgreSQL container is running
docker compose ps postgres
# Test connectivity
docker exec wallow-postgres pg_isready -U wallow
# Check logs
docker compose logs postgres
Solutions
Container not running:
cd docker && docker compose up -d postgres
Wrong connection string:
Check appsettings.Development.json or environment variables:
{
"ConnectionStrings": {
"DefaultConnection": "Host=localhost;Port=5432;Database=wallow;Username=wallow;Password=wallow"
}
}
Database not initialized:
# Recreate with init scripts
cd docker && docker compose down -v
docker compose up -d postgres
PostgreSQL not accepting connections:
FATAL: no pg_hba.conf entry for host
Check that the database user has proper permissions. Module schemas are not created by the API
on startup — Wallow.MigrationService applies them, and the API only migrates inline in the
Testing environment, where Testcontainers hands it an empty database. Under Aspire
(pnpm backend), in the e2e stack and in production, the migration service runs to completion
first and everything else waits on it.
Valkey/Redis Connection Problems
Symptom
StackExchange.Redis.RedisConnectionException: It was not possible to connect to the redis server(s)
Diagnosis
# Check Valkey container
docker compose ps valkey
# Test connectivity
docker exec wallow-valkey valkey-cli ping
# Should return: PONG
# Check logs
docker compose logs valkey
Solutions
Container not running:
cd docker && docker compose up -d valkey
Wrong connection string: the dev stack's Valkey requires a password, so a bare localhost:6379
authenticates as nobody and the connection is refused. appsettings.Development.json ships the
working value:
{
"ConnectionStrings": {
"Redis": "localhost:6379,password=WallowValkey123!,abortConnect=false"
}
}
Memory limit exceeded:
# Check memory usage
docker exec wallow-valkey valkey-cli info memory
# Clear cache if needed
docker exec wallow-valkey valkey-cli FLUSHALL
TLS/SSL configuration issues: For production with TLS:
"Redis": "localhost:6379,ssl=true,abortConnect=false"
API Returns 503 for Everything (First-Run Setup Mode)
Symptom
On a fresh deployment, nearly every request answers 503 Service Unavailable as
application/problem+json whose code is Setup.Required — even though every container is
healthy. The title and detail are the generic 5xx wording; only the code distinguishes the
setup lock from a real outage. Only /v1/identity/setup, /health, /.well-known, /connect, /openapi, and
/scalar respond normally.
Cause
This is not an outage. The production seed deliberately creates no administrator, so
SetupMiddleware (api/src/Wallow.Api/Middleware/SetupMiddleware.cs) locks the API until one
exists.
Solution
Check the status probe, then create the bootstrap admin — via the auth app's setup page or
POST /v1/identity/setup/admin:
curl https://your-domain/api/v1/identity/setup/status
# → {"setupRequired": true}
The 503s stop as soon as the admin exists. See the deployment guide's Setup mode section for the full request.
2. Authentication Issues
JWT Validation Failures
Symptom
Microsoft.IdentityModel.Tokens.SecurityTokenSignatureKeyNotFoundException: IDX10500: Signature validation failed
or
401 Unauthorized
WWW-Authenticate: Bearer error="invalid_token"
Diagnosis
# Check that the API is running and healthy
curl http://localhost:5001/health/ready
Solutions
Wrong authentication configuration:
Check appsettings.json for correct OpenIddict settings.
Clock skew between server and client:
IDX10222: Lifetime validation failed. The token is expired.
Ensure system clocks are synchronized. JWT has a 5-minute tolerance by default.
Token Expiration Problems
Symptom
Token expired at [timestamp]
Solutions
All tokens come from the OpenIddict token endpoint, POST /connect/token. It takes
application/x-www-form-urlencoded parameters, not JSON. There is no email/password token endpoint:
the API supports the authorization code (with PKCE), refresh token, and client credentials grants
only, so a browser user re-authenticates by going back through /connect/authorize.
Get a fresh token (service account / client credentials):
curl -s -X POST http://localhost:5001/connect/token \
-d "grant_type=client_credentials" \
-d "client_id=<your-client-id>" \
-d "client_secret=<your-client-secret>" \
-d "scope=inquiries.read inquiries.write"
Use a refresh token:
curl -s -X POST http://localhost:5001/connect/token \
-d "grant_type=refresh_token" \
-d "refresh_token=YOUR_REFRESH_TOKEN" \
-d "client_id=<your-client-id>" \
-d "client_secret=<your-client-secret>"
Refresh tokens are rolling: issuing a new one revokes the old one, so a retry with an already-used
refresh token fails once the reuse leeway (OpenIddict:RefreshTokenReuseLeewaySeconds, default 30)
has passed — a later replay revokes the whole token family. Sliding expiration is disabled, meaning
the refresh window does not extend on use. Access tokens default to 15 minutes
(OpenIddict:AccessTokenLifetimeMinutes). Refresh-token lifetime is per client: the client's own
refreshTokenLifetime if set, else 7 days for a seeded first-party client, 1 day for any other
application, with OpenIddict:RefreshTokenLifetimeDays (default 7) as the fallback for clients
that carry no per-client value — see the
Configuration guide.
Missing Claims/Permissions
Symptom
403 Forbidden
{
"type": "about:blank",
"title": "Forbidden",
"status": 403,
"detail": "The authenticated identity lacks the permission this resource requires.",
"code": "Auth.Forbidden",
"traceId": "00-abc123def456...-01"
}
Diagnosis
Decode your JWT token at https://jwt.io and check:
roleclaim - Should contain role namesorganizationclaim - Should contain tenant ID
Solutions
User missing role: Assign the required roles via the Identity module's user management API.
Permission not mapped to role:
Check PermissionExpansionMiddleware and role-to-permission mappings in:
api/src/Modules/Identity/Wallow.Identity.Infrastructure/Authorization/PermissionExpansionMiddleware.cs
Organization claim missing: Ensure user belongs to an organization via the Identity module's organization management API.
Tenant Resolution Failures
Symptom
There is no dedicated exception type for this — the observable is ITenantContext.IsResolved
returning false in a handler, and whatever the handler does next when it has no tenant (usually a
null-reference or an empty result set where rows were expected).
Diagnosis
Check ITenantContext.IsResolved in your handler before reading TenantId.
Solutions
Missing organization claim:
The TenantResolutionMiddleware reads tenant from JWT organization claim or X-Tenant-Id header.
Ensure your JWT contains:
{
"organization": "00000000-0000-0000-0000-000000000001"
}
Raw SQL without a tenant filter:
EF Core's tenant query filters do not apply to FromSql/ExecuteSql, so raw SQL must filter by
tenant itself:
await dbContext.Announcements
.FromSql($"SELECT * FROM announcements.announcements WHERE tenant_id = {_tenantContext.TenantId.Value}")
.ToListAsync(cancellationToken);
Test environment:
In tests, WallowApiFactory registers a fixed tenant context. If you need a different tenant, use the test headers:
client.DefaultRequestHeaders.Add("X-Tenant-Id", "your-tenant-guid");
3. Database Issues
Migration Conflicts
Symptom
Microsoft.EntityFrameworkCore.DbUpdateException: An error occurred while saving the entity changes
---> Npgsql.PostgresException: 42P01: relation "identity.users" does not exist
Diagnosis
# Check migration status
dotnet ef migrations list \
--project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
--startup-project api/src/Wallow.Api \
--context IdentityDbContext
Solutions
Apply pending migrations:
dotnet ef database update \
--project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
--startup-project api/src/Wallow.Api \
--context IdentityDbContext
Migration history mismatch:
# Reset database (WARNING: deletes all data)
cd docker && docker compose down -v
docker compose up -d postgres
# Reapply every module's migrations. The API does NOT do this on startup —
# starting it against the empty database just fails differently.
dotnet run --project api/src/Wallow.MigrationService
Conflicting migration:
The migration '20260215_AddNewField' has already been applied to the database
# Remove the conflicting migration
dotnet ef migrations remove \
--project api/src/Modules/Identity/Wallow.Identity.Infrastructure \
--startup-project api/src/Wallow.Api \
--context IdentityDbContext
EF Core Tracking Issues
Symptom
System.InvalidOperationException: The instance of entity type 'Invoice' cannot be tracked because another instance with the same key value is already being tracked
Solutions
Use AsNoTracking for read-only queries:
var invoices = await _context.Invoices
.AsNoTracking()
.Where(i => i.Status == InvoiceStatus.Pending)
.ToListAsync();
Detach existing entity:
var existingEntry = _context.Entry(existingInvoice);
existingEntry.State = EntityState.Detached;
Use new DbContext scope:
using var scope = _serviceProvider.CreateScope();
var context = scope.ServiceProvider.GetRequiredService<IdentityDbContext>();
// Now you have a fresh tracking context
Connection Pool Exhaustion
Symptom
Npgsql.NpgsqlException: The connection pool has been exhausted
---> System.InvalidOperationException: Timeout expired. The timeout period elapsed prior to obtaining a connection from the pool.
Diagnosis
-- Check active connections
SELECT count(*) FROM pg_stat_activity WHERE datname = 'wallow';
-- See connection details
SELECT pid, usename, application_name, state, query_start
FROM pg_stat_activity
WHERE datname = 'wallow'
ORDER BY query_start DESC;
Solutions
Increase pool size:
{
"ConnectionStrings": {
"DefaultConnection": "Host=localhost;...;Maximum Pool Size=100;Connection Idle Lifetime=300"
}
}
Dispose connections properly:
// Use 'using' or 'await using' for DbContext
await using var context = await _contextFactory.CreateDbContextAsync();
Close long-running connections:
-- Terminate idle connections older than 10 minutes
SELECT pg_terminate_backend(pid)
FROM pg_stat_activity
WHERE datname = 'wallow'
AND state = 'idle'
AND query_start < now() - interval '10 minutes';
4. Messaging Issues
Wallow uses Wolverine with in-memory messaging. There is no external message broker.
Messages Not Being Delivered
Symptom
- Events published but handlers never execute
- No errors in logs
Solutions
Handler not discovered:
Wolverine discovers handlers only in the assemblies each enabled module declares through
IWallowModule.HandlerAssemblies. Ensure:
- Handler class is
public static - Method is named
HandleorHandleAsync - The handler's assembly is listed in its module's
HandlerAssemblies(every module declares both its.Applicationand its.Infrastructureassembly, so a handler in either is already covered) - The owning module is enabled — a module switched off in
FeatureManagement:Modules.*contributes no handler assemblies at all, so its messages go unhandled
// Correct handler pattern
public static class MyEventHandler
{
public static async Task HandleAsync(MyEvent @event, IMyService service, CancellationToken ct)
{
// Handle event
}
}
Check Wolverine envelope tables for errored messages:
SELECT * FROM wolverine.wolverine_incoming_envelopes WHERE status = 'error';
Handler Not Being Discovered
Symptom
Wolverine.Runtime.UnknownMessageTypeException: Unknown message type 'MyEvent'
Diagnosis
Check Wolverine's discovered handlers at startup in logs.
Solutions
Ensure handler follows conventions:
// Must be public static class
public static class MyEventHandler
{
// Method must be Handle or HandleAsync
// First parameter must be the message type
public static async Task HandleAsync(MyEvent @event, ILogger<MyEvent> logger)
{
// ...
}
}
Assembly not included in discovery:
Program.cs does not scan for assemblies — it hands Wolverine exactly what the enabled modules
declare:
Assembly[] handlerAssemblies =
[
typeof(Wallow.Api.WallowModules).Assembly, // the host
typeof(IWallowModule).Assembly, // Wallow.Shared.Infrastructure — no module owns it
.. enabledModules.SelectMany(module => module.HandlerAssemblies),
];
So the fix is in the owning module, not in Program.cs. Add the assembly to that module's
HandlerAssemblies in api/src/Modules/{Module}/Wallow.{Module}.Infrastructure/Modules/{Module}Module.cs:
public IReadOnlyList<Assembly> HandlerAssemblies =>
[
typeof(CreateThingHandler).Assembly, // .Application
typeof({Module}Module).Assembly, // .Infrastructure
typeof(MyExtraHandler).Assembly, // the assembly that was missing
];
A brand-new module also needs its entry in WallowModuleRegistry.All
(api/src/Wallow.Modules.Registry/WallowModuleRegistry.cs) — nothing else discovers it.
Outbox Not Processing
Symptom
Messages stuck in wolverine.wolverine_outgoing_envelopes table.
Diagnosis
-- Check outbox status
SELECT status, count(*)
FROM wolverine.wolverine_outgoing_envelopes
GROUP BY status;
-- View stuck messages
SELECT * FROM wolverine.wolverine_outgoing_envelopes
WHERE status = 'scheduled'
ORDER BY scheduled_time;
Solutions
Wolverine agent not running: The Wolverine durability agent processes the outbox. Ensure it's enabled:
opts.PersistMessagesWithPostgresql(connectionString, "wolverine");
Database transaction not committed: Messages are only sent when the transaction commits:
// Ensure SaveChanges is called
await _context.SaveChangesAsync();
// Outbox messages are now ready for sending
Agent polling interval: By default, Wolverine polls every 5 seconds. For debugging:
opts.Durability.PollingInterval = TimeSpan.FromSeconds(1);
5. Test Failures
Testcontainers Not Starting
Symptom
Docker.DotNet.DockerApiException: Docker API responded with status code=InternalServerError
or
Testcontainers.Containers.ContainerStartException: The container did not start in time
Diagnosis
# Verify Docker is running
docker info
# Check Docker resources
docker system info | grep -E "CPUs|Memory"
Solutions
Docker not running: Start Docker Desktop or Docker daemon.
Insufficient resources: In Docker Desktop settings, increase:
- Memory: At least 4GB
- CPUs: At least 2
Port conflicts:
# Find conflicting ports
lsof -i :5432
lsof -i :6379
Container image not found:
# Pull images manually
docker pull postgres:18-alpine
docker pull valkey/valkey:8-alpine
Timeout too short:
// Increase wait time in test fixture
private readonly PostgreSqlContainer _postgres = new PostgreSqlBuilder()
.WithWaitStrategy(Wait.ForUnixContainer()
.UntilPortIsAvailable(5432)
.WithTimeout(TimeSpan.FromMinutes(2)))
.Build();
Parallel Test Conflicts
Symptom
Tests pass individually but fail when run together:
System.InvalidOperationException: Database is in use by another process
Solutions
Use test collection to run sequentially:
[Collection("Database")]
public class InvoiceTests : IClassFixture<WallowApiFactory>
{
// Tests in same collection run sequentially
}
Isolate test data:
// Use unique identifiers per test
var invoiceId = Guid.NewGuid();
var tenantId = Guid.NewGuid();
Disable parallel execution:
In xunit.runner.json:
{
"parallelizeTestCollections": false
}
SignalR Test Issues
Symptom
System.IO.IOException: The server returned status code '401' when status code '101' was expected
Solutions
Include auth token in connection:
var connection = new HubConnectionBuilder()
.WithUrl($"{_factory.Server.BaseAddress}hubs/realtime", options =>
{
options.AccessTokenProvider = () => Task.FromResult<string?>(_token);
options.HttpMessageHandlerFactory = _ => _factory.Server.CreateHandler();
})
.Build();
Use TestAuthHandler headers:
// The TestAuthHandler reads auth from query parameters too
var url = $"{baseUrl}hubs/realtime?access_token={token}";
Wait for connection:
await connection.StartAsync();
// Give SignalR time to establish connection
await Task.Delay(500);
Integration Test Authentication
Symptom
Tests return 401 even with TestAuthHandler.
Solutions
Ensure test scheme is used:
WallowApiFactory configures a "Test" authentication scheme:
services.AddAuthentication("Test")
.AddScheme<AuthenticationSchemeOptions, TestAuthHandler>("Test", options => { });
Set required headers:
client.DefaultRequestHeaders.Add("X-Test-User-Id", Guid.NewGuid().ToString());
client.DefaultRequestHeaders.Add("X-Test-Roles", "admin");
Skip auth for specific tests:
client.DefaultRequestHeaders.Add("X-Test-Auth-Skip", "true");
6. Build Issues
Package Restore Failures
Symptom
error NU1101: Unable to find package Wallow.Storage.Domain
Solutions
Restore from solution root:
cd /path/to/Wallow
dotnet restore
Clear NuGet cache:
dotnet nuget locals all --clear
dotnet restore
Check package sources:
dotnet nuget list source
# Ensure nuget.org is present
Verify network connectivity:
curl https://api.nuget.org/v3/index.json
Project Reference Problems
Symptom
error CS0246: The type or namespace name 'AnnouncementDto' could not be found
Solutions
Check project references:
# View project references
dotnet list api/src/Modules/Announcements/Wallow.Announcements.Api/Wallow.Announcements.Api.csproj reference
Add missing reference:
dotnet add api/src/Modules/Announcements/Wallow.Announcements.Api/Wallow.Announcements.Api.csproj \
reference api/src/Modules/Announcements/Wallow.Announcements.Application/Wallow.Announcements.Application.csproj
Clean and rebuild:
dotnet clean
dotnet build
Assembly Conflicts
Symptom
System.IO.FileLoadException: Could not load file or assembly 'Newtonsoft.Json, Version=13.0.0.0'
Solutions
Check for version conflicts:
All package versions are centrally managed in api/Directory.Packages.props:
grep -r "Newtonsoft.Json" api/Directory.Packages.props
Enable binding redirects: In the project file:
<PropertyGroup>
<AutoGenerateBindingRedirects>true</AutoGenerateBindingRedirects>
</PropertyGroup>
Clear build artifacts:
dotnet clean
rm -rf */bin */obj
dotnet build
7. Frontend Issues
Wallow is a polyglot monorepo: the sections above cover the .NET half, this one covers the pnpm
workspace under apps/ and packages/.
pnpm lint passes but CI fails on a test file
pnpm lint lints source only — it excludes **/*.test.* and **/*.stories.tsx. The excluded
files are linted by the second pass, pnpm lint:tests, which additionally enables oxlint's vitest
plugin. Together the two cover every file exactly once, and pnpm check runs both. Running only
the first one and concluding you are clean is the usual cause.
pnpm lint # source
pnpm lint:tests # test + story files, with the vitest plugin
scripts/lint-tests.sh fails loudly if it enumerates zero files, because oxlint does not expand
globs in path arguments and a silent zero-file pass looks exactly like success.
A change to package source does not take effect
build, typecheck and test run through turbo with content-addressed caching in .turbo/. If
a task replays a stale result, the input you changed is not in that task's hash. Force a run to
confirm:
pnpm build --force # bypass the cache for one run
rm -rf .turbo # or drop the local cache entirely
Caching is local unless TURBO_API/TURBO_TEAM/TURBO_TOKEN are exported — with them set, turbo
also reads and writes the shared remote cache, so a stale entry can come from another machine (see
the developer guide's Turbo remote cache section;
TURBO_REMOTE_CACHE_READ_ONLY=true or unsetting TURBO_TOKEN takes the remote out of the
picture). Note that lint, format, manifest, dependency and export checks are not turbo tasks;
they are root scripts and always run.
Vitest fails to launch a browser
DOM specs run in real headless Chromium through Vitest browser mode with the Playwright provider, never jsdom or happy-dom. A missing browser binary shows up as a launch failure before any test runs:
pnpm exec playwright install chromium
If a spec that touches the DOM instead runs in Node — document is not defined, or
assertBrowserModeSmoke failing — the file is in the wrong vitest project. Check the consumer's
vite.config.ts project inventory; see packages/testing/CLAUDE.md for the split.
E2E suites fail against a backend that is not there
The Playwright suites in apps/wallow-auth/e2e/, apps/wallow-web/e2e/ and
apps/wallow-web/e2e-cross-app/ are backend-dependent. Run them through the one-command runner,
which brings the containerised stack up and tears it down again:
./scripts/e2e.sh
Running pnpm exec playwright test directly against a stack you have not started fails at the
first navigation.
8. Debugging Tips
Enabling Detailed Logging
Serilog configuration:
{
"Serilog": {
"MinimumLevel": {
"Default": "Debug",
"Override": {
"Microsoft": "Information",
"Microsoft.EntityFrameworkCore": "Warning",
"Wolverine": "Debug"
}
}
}
}
EF Core query logging:
{
"Logging": {
"LogLevel": {
"Microsoft.EntityFrameworkCore.Database.Command": "Information"
}
}
}
Wolverine message logging:
Already configured in Program.cs:
opts.ConfigureMessageLogging(); // Logs message execution
Checking Wolverine Envelope Tables
View pending outbox messages:
SELECT
id,
message_type,
destination,
status,
scheduled_time,
attempts
FROM wolverine.wolverine_outgoing_envelopes
ORDER BY scheduled_time DESC
LIMIT 50;
View incoming (inbox) messages:
SELECT
id,
message_type,
status,
received_at
FROM wolverine.wolverine_incoming_envelopes
ORDER BY received_at DESC
LIMIT 50;
Clear stuck messages:
-- WARNING: This may lose messages
DELETE FROM wolverine.wolverine_outgoing_envelopes WHERE status = 'error';
Quick Diagnostic Commands
Check all services:
cd docker && docker compose ps
docker compose logs --tail=50
Test database connectivity:
docker exec wallow-postgres psql -U wallow -d wallow -c "SELECT 1"
Test Redis/Valkey connectivity:
docker exec wallow-valkey valkey-cli ping
View application logs:
# If running via dotnet run
# Logs output to console
# If running in Docker
docker logs wallow-api --tail=100 -f
Reset everything:
cd docker && docker compose down -v && docker compose up -d
Quick Reference: Error Code Mapping
| HTTP Status | Exception Type | Meaning |
|---|---|---|
| 400 | ValidationException, ArgumentException |
Invalid request data |
| 401 | UnauthorizedAccessException |
Authentication failed |
| 403 | ForbiddenAccessException |
Missing permission |
| 404 | EntityNotFoundException |
Resource not found |
| 422 | BusinessRuleException |
Business rule violation |
| 500 | Unhandled exception | Server error |
Getting Help
If you cannot resolve an issue:
Check existing documentation:
docs/getting-started/developer-guide.md- Development setupdocs/operations/deployment.md- Production deployment
Search the codebase:
grep -r "error message" api/src/Check recent commits:
git log --oneline -20Run architecture tests:
./scripts/run-tests.sh arch