Running four client VPNs at once when their networks collide
Two of the companies whose production I look after both use
10.20.0.0/16 internally. Same address space, completely different
servers behind it. Neither of them did anything wrong: it is a private range, and
picking it is the obvious thing to do.
It stops being obvious on the machine in the middle. I need both at the same
time, and one routing table cannot hold two different meanings for
10.20.1.5.
I worked around this for months before fixing it properly. The workarounds are worth describing, because each one looks reasonable until you see what it actually does.
The three answers that do not work
Disconnect one to use the other
This is what almost everybody does, and it is what I did. It works, in the sense that nothing is ever wrong at any single moment. But it quietly shapes how you work: you stop checking a second client's dashboard while you are in the middle of something, you batch work by client because switching costs a minute, and when something breaks for client B while you are inside client A's network, you are not one command away from looking, you are one reconnect away.
Let the routing metrics sort it out
They do not, and the way they fail is genuinely nasty. One of the clients uses
GlobalProtect. When that tunnel is up it publishes routes for essentially every
RFC1918 10.x.0.0/16 there is, at metric 0.
Every other client's tunnel has its own /16 sitting at metric 1.
Lower metric wins. So connecting one client silently takes away every other
client's network, while both tunnels continue to report themselves as happily
connected. Nothing logs an error. The other VPN's status is green. The packets
simply go somewhere else.
Check which interface the OS actually picked, not whether the tunnel says it
is up. On Windows, Find-NetRoute -RemoteIPAddress <target>
names the adapter that will carry the packet. If it comes back as the
GlobalProtect adapter while you are trying to reach a different client, that is
the whole bug.
Add a more specific route
A /24 beats a /16 regardless of metric, so a manual
route does win. It also needs administrator rights, has to be re-added every time
the tunnel reconnects, and leaves you maintaining a private list of which client
owns which subnet. It is a patch you have to keep applying, which means one day
you will not.
The actual fix: one network namespace per client
The reason all of this is painful is that there is one routing table. A Linux network namespace gives you another one. Each namespace has its own interfaces, its own routes, its own iptables rules, and no opinion whatsoever about what any other namespace is doing.
So: one namespace per client, and the VPN client process runs
inside it. Its tun device and all the routes it installs are
created in there and stay in there. Two clients announcing the same
10.20.0.0/16 stop being a conflict, because the two routes live in
two different tables that never meet.
The shape of it, with the error handling stripped out:
# one namespace per client
ip netns add vpn-acme
# a veth pair: one end stays on the host, the other goes inside
ip link add veth-acme type veth peer name veth-acme-ns
ip link set veth-acme-ns netns vpn-acme
# a tiny /30 for the pair, unique per client
ip addr add 10.201.7.1/30 dev veth-acme
ip link set veth-acme up
ip netns exec vpn-acme ip addr add 10.201.7.2/30 dev veth-acme-ns
ip netns exec vpn-acme ip link set veth-acme-ns up
ip netns exec vpn-acme ip link set lo up
ip netns exec vpn-acme ip route add default via 10.201.7.1
# let the namespace reach the internet to build the tunnel
iptables -t nat -A POSTROUTING -s 10.201.7.0/30 -o eth0 -j MASQUERADE
# and the VPN itself, inside
ip netns exec vpn-acme openconnect --protocol=anyconnect vpn.acme.example
Keep a fixed number per client for that /30 so two clients can
never be handed the same one. A five line table in the script is enough.
Reaching things, without thinking about it
A namespace you have to remember to enter is a namespace you will forget to enter, and the failure mode is confusing: you will SSH to a valid address and reach the wrong company's server. So the entry point is a wrapper:
cx acme ssh [email protected]
cx acme curl -s http://10.20.1.5:9090/api/v1/targets
It runs the command inside that client's namespace, as you, with your environment. There is no mode to be in and nothing to remember to turn off. If the client name is wrong the command fails immediately instead of succeeding against somebody else's network.
The part I did not expect: SSH host keys
The first time I had two tunnels up at once, SSH started refusing connections with host key warnings. It was right to.
10.20.1.5 is a different physical machine depending on whose
network you are inside. One known_hosts file cannot represent that:
it maps an address to a key, and here one address honestly has several keys. The
warning was not noise, it was the tool correctly reporting that the thing behind
the address had changed.
So each client gets its own file. The wrapper exports the client name and
~/.ssh/config uses it:
Match exec "test -n \"$VPN_CLIENT\""
UserKnownHostsFile ~/.ssh/known_hosts.%{VPN_CLIENT}
IdentityFile ~/ops/%{VPN_CLIENT}/id_%{VPN_CLIENT}
Which has a side effect I like more than the fix itself: host key checking becomes meaningful again. A changed key inside one client's file is now a real signal rather than the thing you dismiss twenty times a day.
Two traps worth knowing before you build this
DNS inside a namespace is not what you expect
Namespaces isolate interfaces and routes. They do not isolate
/etc/resolv.conf, which is just a file. On a normal machine you can
mount a different one per namespace. Under WSL you cannot usefully do that,
because /etc/resolv.conf is a symlink into a shared mount, so there
is exactly one for everything.
Two consequences. Resolving the VPN gateway's hostname from inside the namespace fails, which you fix by resolving it on the host first and passing the answer in:
openconnect --resolve vpn.acme.example:203.0.113.10 vpn.acme.example
And more importantly, the VPN's own connect script will happily rewrite that
shared file with the client's internal DNS servers, breaking name resolution for
the entire machine including every other client. Clear
INTERNAL_IP4_DNS in a wrapper script before the real
vpnc-script runs, and address internal hosts by IP.
This is a Linux-side fix only
The namespace isolates traffic from processes you launch inside it. If you are on Windows with WSL, applications on the Windows side are not affected and Windows' own routing table is still a free for all. For a workflow that is entirely SSH, curl and tooling from inside Linux, that is fine. For anything that has to go through a desktop application, it is not, and you are back to reconnecting.
Where it ended up
Four clients, four tunnels, all up at the same time, overlapping ranges irrelevant. On a server it runs unattended as one systemd unit per client, with the password read from a root-only file and a keepalive ping every sixty seconds because several gateways drop idle tunnels without saying so.
What I would tell a version of me from a year ago: the reconnect dance does not feel like a problem because each individual reconnect is only a minute. The cost is not the minute. The cost is that you stop looking at things.