Rıza KorkusuzTechnical Support
0 / 13 reviewed
← Back to Portfolio
Study roadmap · Computer networks

Networking, from the cable to the cloud.

Thirteen modules that build on each other, written from a support point of view: what each layer does, how it breaks, which commands prove it, and how to explain it in an interview.

13 modules~82 min total21 animated diagramsWindows · macOS · Linux commands48 self-check questions
Read bottom-upModules 1–5 follow the stack from the wire to transport. Most tickets get solved there. Modules 6–9 cover DNS, the web, and proxies.
Run every commandEach module has a symptom → cause → check table and the same commands for Windows, macOS and Linux.
Test yourselfAnswer the "Check yourself" cards out loud, then mark the module reviewed. Two clear minutes per topic is the interview bar.
1

What a network really is

MODULE 01 · ~6 min read

Devices following shared rules to move data in small, addressed pieces, and a way to think about it that survives any ticket.

A network is two or more devices that exchange data by following the same rules. That is the whole idea. A laptop, a badge reader, a printer and a cloud server can all talk because each one implements the same protocols: agreements about how data is formatted, addressed, sent, acknowledged and checked.

Why start here if you are going into technical support? Because "the internet is down" is never the real problem. It is a symptom. The people who fix tickets quickly have a mental model of the pieces involved, so they can ask "which piece is failing?" instead of rebooting things at random. This module builds that model.

Core ideas

Protocols are layered agreements

No single protocol does everything. Opening a web page uses several at once, each solving one problem:

  • HTTP decides what the browser asks for and what the server sends back.
  • TCP makes sure those bytes arrive complete and in order.
  • IP gets each piece from your network to the server's network, across many routers.
  • Ethernet or Wi-Fi moves the piece across the one local link it is on right now.

Each protocol trusts the one below it to do its job and ignores the details. HTTP has no idea whether you are on Wi-Fi or a cable. That separation is why you can troubleshoot one layer at a time.

Packets, not streams

Data is cut into packets. On a normal Ethernet or Wi-Fi network a packet carries at most about 1,500 bytes, so a 3 MB photo becomes roughly two thousand packets. Each packet carries its own addressing, travels independently, and is reassembled at the far end.

This is called packet switching. Many conversations share the same links, taking turns packet by packet. If one packet is lost, only that packet is resent. The old telephone network did the opposite (a dedicated circuit for each call), which is why it wasted capacity whenever nobody was speaking.

Encapsulation: envelopes inside envelopes

As data goes down the stack on the sender, each layer adds its own header in front of what it received:

[ Ethernet hdr | IP hdr | TCP hdr | HTTP data ........ | FCS ]
  MAC → MAC     IP → IP  port → port  "GET /index.html"    checksum

The receiver peels the layers off in reverse. One detail worth remembering for interviews: at every router along the way, the outer Ethernet header is removed and a new one is written for the next link. The MAC addresses change at every hop; the IP addresses stay the same end to end (unless NAT rewrites them, covered in module 2).

Two models: OSI vocabulary, TCP/IP reality

The OSI model has seven layers: 1 Physical, 2 Data Link, 3 Network, 4 Transport, 5 Session, 6 Presentation, 7 Application. Hardly any software is built in those exact seven pieces, but the numbers are used as shorthand everywhere: "layer 2 issue" means switching/VLANs, "layer 3" means IP and routing, "layer 7 load balancer" means it understands HTTP.

The TCP/IP model is what actually runs: Link, Internet, Transport, Application. Layers 5–7 of OSI are folded into "application". Know the OSI numbers so you can speak the language; reason with the four TCP/IP layers.

The building blocks you will touch

  • Endpoints / hosts: laptops, phones, servers, printers, IP phones, cameras.
  • Switch: connects devices inside one local network (layer 2).
  • Router: connects different networks together and picks paths between them (layer 3). A home "router" is usually a router, switch, Wi-Fi access point, firewall and DHCP server in one box.
  • Access point: bridges Wi-Fi clients onto the wired network.
  • Firewall: allows or blocks traffic based on rules.
  • Servers and services: DHCP, DNS, web, file, identity.

Bits versus bytes

Network speeds are quoted in bits per second (Mbps). File sizes and download meters use bytes (MB/s). There are 8 bits in a byte, so a 100 Mbps connection tops out around 12 MB/s, and a bit lower after protocol overhead. A user who says "I pay for 100 and only get 11" is usually getting exactly what they pay for.

Worked example

What happens when you open https://intranet.example.com

  1. Link + address: the laptop is already on Wi-Fi and got 10.10.20.57/24, a gateway and DNS servers from DHCP.
  2. DNS: it asks its DNS server for intranet.example.com and gets 10.20.0.15.
  3. Local or remote? 10.20.0.15 is not in 10.10.20.0/24, so the packet must go to the gateway. ARP finds the gateway's MAC address.
  4. TCP: a three-way handshake to port 443 opens a connection.
  5. TLS: the server presents a certificate; the browser checks it and both sides agree on keys.
  6. HTTP: the browser sends GET /; the server answers 200 OK with HTML, and the browser repeats the process for images, scripts and styles.

Every step is a place where things break. The later modules take them one at a time.

How it breaks, and what to check

SymptomLikely causeWhat to check
No link light, Wi-Fi says "no networks"physical layer: cable, port, radio off, airplane modeTry a known-good cable and port; check the Wi-Fi switch/toggle
Address starts with 169.254device never got a DHCP leaseModule 2: DHCP, VLAN, Wi-Fi authentication
Can ping 1.1.1.1 but no website loads by nameDNSModule 6: nslookup against a second resolver
Every site works except onethat service, its DNS, a proxy rule, or a certificateTry it from another network; read the exact error
Everything works, just slowlyperformance: Wi-Fi, saturation, VPN pathModule 12: measure latency, loss and throughput

Commands

TaskWindowsmacOSLinux
Show my IP setupipconfig /allifconfig or System Settings → Networkip addr
Can I reach it?ping 1.1.1.1ping 1.1.1.1ping 1.1.1.1
Which path does it take?tracert 1.1.1.1traceroute 1.1.1.1traceroute 1.1.1.1
Does the name resolve?nslookup example.comnslookup example.comdig example.com
At the help deskScope before you fix. Ask: one user or many? One site or everything? Wired, Wi-Fi or VPN? When did it start, and what changed? Then test bottom-up: link → IP address → gateway → internet by IP → DNS → the application. Write down each result. A ticket that says "gateway pings, 1.1.1.1 pings, DNS fails against 10.10.0.2 but works against 1.1.1.1" gets fixed in minutes.

Check yourself

Answer out loud first, then open the card.

At a router hop, which addresses change: MAC or IP?

The MAC addresses (the outer Ethernet header is rewritten for each link). The source and destination IP stay the same, unless a NAT device rewrites them.

A user on a 200 Mbps plan sees downloads at 23 MB/s. Problem?

No. 200 Mbps ÷ 8 = 25 MB/s, minus protocol overhead. That is normal.

Name the protocols involved in loading an HTTPS page, bottom to top.

Ethernet or Wi-Fi, IP (with ARP to find the gateway), TCP, TLS, HTTP, plus DNS beforehand to find the address.

2

Addresses and boundaries: MAC, IP, DHCP, NAT

MODULE 02 · ~7 min read

Who a device is on the local wire, where it sits in the wider network, how it gets its settings, and where the edges are.

Every device on a network needs a few pieces of information before it can do anything useful: an IP address, a subnet mask, a default gateway and DNS servers. Most of the time DHCP hands these out silently. When it doesn't, the user sees "no internet", and you see a 169.254 address.

This module covers the two kinds of address (hardware MAC and logical IP), the private and special ranges you will see in tickets, how DHCP and NAT work, and the LAN/WAN boundary. If you can read an ipconfig /all output and explain every line, you are ahead of most candidates.

Core ideas

MAC addresses: the local name tag

A MAC address is 48 bits, written as six hex pairs: 3c:22:fb:9a:10:4e (Windows shows it as 3C-22-FB-9A-10-4E). The first half traditionally identifies the manufacturer. A MAC is only meaningful on the local network segment. Routers do not forward it beyond their own link.

Modern phones and Windows/macOS laptops can use a random (private) MAC per Wi-Fi network. It protects privacy but breaks anything keyed to a fixed MAC: DHCP reservations, MAC allowlists, some captive portals and NAC policies. If the second hex digit is 2, 6, A or E (for example a6:… or da:…), it is a locally assigned, likely randomized address.

IPv4 addresses and the subnet mask

An IPv4 address is 32 bits written as four numbers from 0 to 255: 10.10.20.57. The subnet mask splits it into a network part and a host part. /24 (= 255.255.255.0) means the first 24 bits are the network, so 10.10.20.57/24 lives on network 10.10.20.0, with hosts .1 to .254 and broadcast .255.

The device uses the mask for one decision: is the destination on my network (deliver directly) or not (send to the gateway)? Module 4 does the subnet math.

Ranges worth recognizing on sight

RangeMeaning
10.0.0.0/8, 172.16.0.0/12, 192.168.0.0/16Private (RFC 1918). Used inside homes and companies; not routed on the internet.
169.254.0.0/16Link-local / APIPA. The device gave itself an address because DHCP did not answer.
127.0.0.0/8Loopback, "this machine". 127.0.0.1 = localhost.
100.64.0.0/10Carrier-grade NAT. Your ISP is sharing public addresses between customers.
192.0.2.0/24, 198.51.100.0/24, 203.0.113.0/24Reserved for documentation and examples (used on this page).

DHCP: four messages to get online

A new device has no address, so it broadcasts a Discover. A DHCP server answers with an Offer, the client sends a Request for that address, and the server confirms with an Acknowledge, often remembered as DORA. The ACK includes the lease length and options such as router (gateway), DNS servers and domain name.

Leases expire. A client tries to renew at about half the lease time, so a laptop that is on for days keeps the same address. Two things you will meet in offices:

  • DHCP relay (often called an "IP helper"): DHCP Discover is a broadcast and stops at the router, so each VLAN's router interface forwards it to a central DHCP server. A missing relay on a new VLAN = every device gets 169.254.
  • Reservations: the server always gives a given MAC the same IP, used for printers and some servers. Random MACs defeat them.

Default gateway, LAN and WAN

The default gateway is the router interface on your subnet, usually .1 or .254. Anything not local is handed to it. A wrong gateway produces a classic symptom: local printers and file shares work, nothing else does.

The LAN is the network you control: an office floor, a home, a campus. The WAN connects sites to each other and to the internet through ISPs, leased lines, SD-WAN or VPN tunnels. A WLAN is the Wi-Fi part of the LAN. The router at the edge is where LAN meets WAN, and usually where NAT and the firewall live.

NAT and PAT: many devices, one public address

Private addresses can't be used on the internet, so the edge router rewrites them. With PAT (port address translation, what home routers and most offices do), the router replaces the source 10.10.20.57:51544 with its public 203.0.113.7:40001, remembers the mapping, and reverses it when the reply arrives. Hundreds of devices share one public IP this way.

Consequences you will see: inbound connections need an explicit port forward; a "what is my IP" site shows the router's public address, not the laptop's; and with CGNAT (WAN address in 100.64.0.0/10) the ISP is doing another layer of NAT, so port forwarding at home cannot work.

IPv6 in two minutes

IPv6 addresses are 128 bits in hex: 2001:db8:4a1:20::57 (2001:db8::/32 is the documentation range). :: compresses a run of zeros. Every interface has a link-local fe80:: address, and most also get a global address automatically (SLAAC) or from DHCPv6. There is no NAT in typical IPv6 deployments and no ARP; Neighbor Discovery does that job.

Many networks are dual stack. A site that fails only over IPv6 while IPv4 works is a real (if uncommon) ticket. Comparing ping -4 and ping -6 separates them.

Worked example

Reading ipconfig /all line by line

Wireless LAN adapter Wi-Fi:
   Physical Address. . . . . . . . . : 3C-22-FB-9A-10-4E   <- MAC
   DHCP Enabled. . . . . . . . . . . : Yes
   IPv4 Address. . . . . . . . . . . : 10.10.20.57(Preferred)
   Subnet Mask . . . . . . . . . . . : 255.255.255.0       <- /24
   Lease Obtained. . . . . . . . . . : Monday 08:02
   Lease Expires . . . . . . . . . . : Monday 16:02        <- 8 h lease
   Default Gateway . . . . . . . . . : 10.10.20.1
   DHCP Server . . . . . . . . . . . : 10.10.0.5           <- via relay
   DNS Servers . . . . . . . . . . . : 10.10.0.2
                                       10.10.0.3

Healthy. Now picture the same output with IPv4 Address: 169.254.83.12, no gateway, no DHCP server. That laptop never heard back from DHCP: look at Wi-Fi authentication, the switch port's VLAN, the relay, or a full DHCP pool. DNS is irrelevant at this point.

How it breaks, and what to check

SymptomLikely causeWhat to check
169.254.x.x addressno DHCP reply: auth failed, wrong VLAN, relay missing, pool full, server downipconfig /release + /renew; check other devices on same port/SSID; ask network team about pool
"Windows has detected an IP address conflict"two devices with same IP, often a static IP inside the DHCP poolarp -a to see the MAC claiming it; move statics outside the pool or reserve
Local shares work, internet doesn'twrong or unreachable gatewayping the gateway; compare gateway with a working neighbor
Printer's reservation stopped workingdevice using a random MAC, or NIC replacedCompare the MAC on the device with the reservation
Port forward at home does nothingCGNAT, or double NAT behind ISP routerRouter WAN IP in 100.64.0.0/10 or private = CGNAT/double NAT
Guest Wi-Fi full at a big meetingDHCP pool exhausted (short pool, long lease)DHCP server lease table; shorter leases on guest networks

Commands

TaskWindowsmacOSLinux
Full address infoipconfig /allifconfig en0; ipconfig getpacket en0ip addr; nmcli dev show
Renew DHCP leaseipconfig /release then ipconfig /renewsudo ipconfig set en0 DHCPsudo nmcli con up "<name>" or sudo dhclient -r; sudo dhclient
Default gatewayroute print / ipconfignetstat -rn | grep defaultip route
MAC addressgetmac /vifconfig en0 etherip link
My public IPnslookup myip.opendns.com resolver1.opendns.comsame commanddig +short myip.opendns.com @resolver1.opendns.com
At the help deskLook at the address first, every time. 169.254 = DHCP problem, not DNS or "the internet". A correct-looking address in the wrong range (staff laptop in the guest subnet) = VLAN or SSID mapping. Correct range but nothing beyond the LAN = gateway or upstream. On Wi-Fi, "connected, no internet" with a valid address often means a captive portal or DNS, not addressing.

Check yourself

Answer out loud first, then open the card.

A laptop shows 169.254.10.4. Should you flush DNS?

No. It never got a DHCP lease, so it has no gateway or DNS server at all. Fix link/auth/VLAN/DHCP first.

What are the four DHCP messages?

Discover, Offer, Request, Acknowledge (DORA).

Why can 300 office laptops share one public IP?

PAT: the edge router tracks each connection by port and rewrites private source addresses and ports to its single public address.

What does a WAN address of 100.64.8.20 on a home router tell you?

The ISP uses carrier-grade NAT, so inbound port forwards won't work without an ISP change or a tunnel/relay service.

3

Link layer: frames, switches, ARP, VLANs and Wi-Fi

MODULE 03 · ~7 min read

How data crosses one local network, and the cabling, switching and wireless problems that cause a big share of deskside tickets.

Layer 2 is the local hop: from a laptop to the switch, from the switch to the router, from a phone to the access point. Data here travels in frames addressed by MAC. Nothing at this layer knows about the internet. It only knows "which device on this segment gets this frame?"

For deskside support this is where a surprising amount of real work happens: bad cables, wrong wall-port patching, VLAN mistakes, PoE phones, Wi-Fi signal and authentication. These are physical, checkable things, and the first thing an experienced tech looks at.

Core ideas

What is inside an Ethernet frame

| dest MAC | source MAC | [802.1Q tag] | EtherType | payload (≤1500 B) | FCS |
                                          0x0800 = IPv4, 0x86DD = IPv6, 0x0806 = ARP

The FCS is a checksum. If a frame arrives damaged, the receiver discards it silently. On a switch this shows up as CRC / input errors on the port, a strong hint of a bad cable, bad patch, or electrical interference.

How a switch decides where frames go

A switch reads the source MAC of every frame and records "this MAC is behind port 7" in its MAC address table (entries age out after a few minutes of silence). It then looks up the destination MAC:

  • Known → send out that one port only.
  • Unknown, or broadcast ff:ff:ff:ff:ff:ff → flood out every port in the same VLAN.

Old hubs repeated every frame to every port; switches don't, which is why you can't capture a colleague's traffic just by plugging into the same switch. Everything in one VLAN is one broadcast domain: broadcasts reach every device in it, and routers stop them.

ARP: finding a MAC for an IP

IP knows the destination IP; Ethernet needs a MAC. ARP bridges the gap. The laptop broadcasts "who has 10.0.0.1? tell 10.0.0.5". Only the owner answers, directly. The result is cached for a short time (see it with arp -a).

Two practical points. For anything off-subnet, the laptop ARPs for its gateway, not the remote server. And because ARP trusts whoever answers, a rogue device can claim the gateway's IP (ARP spoofing). Managed switches have protections like Dynamic ARP Inspection. IPv6 replaces ARP with Neighbor Discovery.

VLANs, access ports and trunks

A VLAN splits one physical switch into separate logical networks, each with its own subnet: for example VLAN 10 staff, 20 voice, 30 guest, 40 printers. Devices in different VLANs can only talk through a router or layer 3 switch, where firewall rules can be applied.

  • An access port belongs to one VLAN; the PC has no idea VLANs exist.
  • A trunk carries many VLANs between switches, access points and routers, marking each frame with an 802.1Q tag (VLAN ID 1–4094).
  • A desk IP phone usually sits on a port with a voice VLAN: the phone tags its own traffic into voice, and the PC plugged into the phone's back port stays on the data VLAN.

Loops and spanning tree

Connect two switch ports together with a cable (or a user plugs both ends of a patch cable into the wall) and broadcasts circulate forever: a broadcast storm that can freeze a whole floor. Spanning Tree Protocol (STP) detects loops and blocks redundant ports. Classic STP makes a newly connected port wait around 30 seconds before forwarding, which is why switches set end-user ports to "PortFast" / edge mode. Without it, a PC can boot faster than the port comes up and get a 169.254 address.

Cabling and physical checks

  • Copper Ethernet runs are limited to 100 m. Cat5e handles 1 Gbps; Cat6/6A for higher speeds.
  • Gigabit uses all four pairs; 100 Mbps uses two. A cable with one damaged pair often links at 100 Mbps instead of 1 Gbps: a great clue when "the network is slow" for one desk.
  • Speed/duplex should auto-negotiate on both ends. A mismatch (one side forced) causes errors and terrible throughput while ping still looks fine.
  • PoE powers phones, access points and cameras over the cable. A phone that boots in one port but not another may be on a non-PoE port, or the switch has run out of PoE budget.

Wi-Fi essentials

Wi-Fi (802.11) is layer 1–2 over radio. The SSID is the network name; each access point radio has its own BSSID (a MAC). Clients choose and roam between APs themselves, which is why one laptop can "stick" to a distant AP.

  • Bands: 2.4 GHz travels further through walls but is crowded (only channels 1, 6, 11 don't overlap); 5 GHz and 6 GHz (Wi-Fi 6E/7) are faster with shorter range.
  • Signal (RSSI) is in negative dBm: about −50 is excellent, −67 is a common minimum target for voice and video, below −75 expect trouble. Noise and interference matter as much as raw signal.
  • Security: WPA2/WPA3-Personal uses a shared passphrase; Enterprise uses 802.1X, where each user or device authenticates (password or certificate) to a RADIUS server. Enterprise Wi-Fi failures are often credentials, expired certificates, or a device not trusting the RADIUS server's certificate.
  • Captive portals (hotels, guest Wi-Fi) give you an address but block traffic until you accept terms in a browser.
Worked example

Desk 4B gets the wrong address

A new hire at desk 4B gets 10.30.8.112. Staff laptops should be in 10.10.x.x; 10.30 is the guest VLAN. The laptop is fine, DHCP is fine: it's doing exactly what the port told it to. Steps:

  1. Note the wall port label (e.g. 4B-02).
  2. Try the laptop on a neighbor's working port: it gets 10.10.x.x, so the laptop is ruled out.
  3. Raise it with the network team: "Wall port 4B-02 appears to be on VLAN 30 (guest). Neighboring 4B-01 is on staff. Please check the patch and switchport VLAN."

That's a five-minute ticket with a precise handoff, instead of a reimage.

How it breaks, and what to check

SymptomLikely causeWhat to check
No link lightcable unplugged or broken, port disabled, NIC disabledKnown-good cable; another port; NIC enabled in OS
Links at 100 Mbps, slow transfersdamaged pair in cable or bad patchCheck link speed; replace patch cable; ask for port error counters
Right SSID/port, wrong subnetport or SSID mapped to wrong VLANCompare with neighbor; give port label to network team
Whole floor slow, lights blinking wildlyswitching loop / broadcast stormLook for a cable looped between two wall ports; escalate immediately
IP phone works, PC behind it gets nothingPC port/data VLAN misconfigured, or phone's PC port disabledPlug PC direct to wall to compare
Enterprise Wi-Fi keeps asking for passwordexpired password or cert, wrong EAP settings, server cert not trustedEvent log / Wi-Fi report; re-enroll certificate; forget and re-add network
Wi-Fi drops in one meeting roomweak signal, interference, sticky roamingCheck RSSI there; compare 2.4 vs 5 GHz; report AP coverage

Commands

TaskWindowsmacOSLinux
ARP / neighbor cachearp -aarp -aip neigh
Clear ARP cachearp -d * (admin)sudo arp -a -dsudo ip neigh flush all
Link speedGet-NetAdapter | ft Name,Status,LinkSpeedifconfig en0 | grep mediaethtool eth0
Wi-Fi signal and bandnetsh wlan show interfacesOption-click the Wi-Fi icon; system_profiler SPAirPortDataTypenmcli dev wifi; iw dev wlan0 link
Wi-Fi history reportnetsh wlan show wlanreportWireless Diagnostics appjournalctl -u NetworkManager
At the help deskCheck the physical layer first, it's quick: link lights, a known-good patch cable, link speed, and the wall port label. Swap one variable at a time (cable, then port, then laptop) so you know which change fixed it. On Wi-Fi, record SSID, band, signal (dBm) and the location; "Wi-Fi is bad" without those can't be acted on.

Check yourself

Answer out loud first, then open the card.

What does a switch do with a frame for a MAC it hasn't learned yet?

Floods it out every port in that VLAN (except the one it came in on), then learns from the reply.

Why might a PC link at 100 Mbps on a gigabit port?

Gigabit needs all four pairs; a damaged pair or poor patch falls back to 100 Mbps, which only uses two.

A laptop needs to reach 8.8.8.8. Whose MAC does it ARP for?

Its default gateway's. Off-subnet traffic is always framed to the gateway.

Difference between WPA2-Personal and WPA2-Enterprise?

Personal: one shared passphrase. Enterprise: each user/device authenticates via 802.1X to a RADIUS server.

4

Network layer: IP routing, subnets and ICMP

MODULE 04 · ~6 min read

How packets cross from one network to another, how to do subnet math in your head, and what ping and traceroute actually measure.

Layer 3 moves packets between networks. Your laptop doesn't know the path to a server in another city. It only knows "local, or give it to the gateway". Each router repeats that small decision with its own table, hop by hop, until the packet arrives.

For support, this layer explains the confusing tickets: "I can reach some things but not others", "it works in the office but not on VPN", "ping fails but the website loads". Subnet math also comes up in interviews, so we'll make it quick.

Core ideas

What an IP packet carries

The IP header holds source and destination addresses, a TTL (time to live, called hop limit in IPv6), the protocol inside (TCP = 6, UDP = 17, ICMP = 1), and fragmentation flags. IP is best effort: no guarantee of delivery or order. Reliability, if needed, comes from TCP above it.

Each router decrements TTL by one. If it reaches zero the packet is dropped and an ICMP "time exceeded" goes back to the sender. This stops packets looping forever, and traceroute is built on it.

Local or remote: the first decision

A host compares the destination with its own network using the mask. Example: laptop 10.1.4.20/22.

  • A /22 mask is 255.255.252.0. The third-octet block size is 256 − 252 = 4, so networks start at 0, 4, 8, 12…
  • 20's network is 10.1.4.0/22, covering 10.1.4.0 – 10.1.7.255.
  • Destination 10.1.6.9: local, ARP for it directly. Destination 10.1.8.1: remote, send to the gateway.

If someone typed /24 instead, the laptop would think 10.1.6.9 is remote and send it to the gateway. Usually still works, oddly, but not always, which is what makes mask errors confusing.

Subnet math you can do in your head

PrefixMaskAddressesUsable hosts
/30255.255.255.25242 (router links)
/29255.255.255.24886
/28255.255.255.2401614
/27255.255.255.2243230
/26255.255.255.1926462
/25255.255.255.128128126
/24255.255.255.0256254
/16255.255.0.065,53665,534

Usable = total − 2 (the network address and the broadcast address). To find which block an address is in: block size = 256 − the interesting mask octet, then round down to a multiple of it. 192.168.5.77/26 → block 64 → network .64, broadcast .127, hosts .65–.126.

Routing tables and longest match

Every router (and your laptop) has a routing table: destination network → next hop or interface. Entries come from directly connected networks, static routes, and dynamic routing protocols. The default route 0.0.0.0/0 matches everything and is used when nothing more specific does.

When several routes match, the most specific (longest prefix) wins. A packet for 10.8.3.4 matches 0.0.0.0/0, 10.0.0.0/8 and 10.8.0.0/16; it goes via the /16. That rule is exactly how VPN clients steer company traffic into the tunnel while the rest goes out locally.

Inside organizations, routers share routes with protocols like OSPF; between ISPs, BGP carries the internet's routes. You won't configure them in a support role, but "a BGP issue at the provider" is a real explanation for a large regional outage.

ICMP: the network's error messages

  • Echo request / reply (type 8 / 0): ping.
  • Destination unreachable (type 3): with codes such as network unreachable, host unreachable, port unreachable, and fragmentation needed.
  • Time exceeded (type 11): TTL hit zero; traceroute uses it.

Firewalls often drop ping, so no reply does not prove a host is down. Blocking all ICMP is a mistake: "fragmentation needed" messages drive Path MTU Discovery, and without them some connections hang on large packets (more in module 12).

VPNs: routing in disguise

When a VPN connects, it adds a virtual interface and routes. Full tunnel sends everything through the company. Split tunnel sends only company networks (say 10.0.0.0/8) through the tunnel and everything else directly.

A classic failure: the user's home network is 192.168.1.0/24 and the office file server is also in 192.168.1.0/24. The laptop treats the server as local and never sends it into the tunnel. Fix: company networks should avoid the common home ranges, or the VPN must push a more specific route.

Worked example

Reading a traceroute

A user in the office reports a slow SaaS app at 198.51.100.20. The terminal on the right shows the trace:

  • Hop 1 is the office gateway (2 ms). Hop 2 is in 100.64.x.x: the ISP.
  • Hop 3 shows * * *: that router doesn't answer traceroute probes. Later hops respond, so it is forwarding fine. Not a problem.
  • Latency jumps from ~14 ms at hop 4 to ~95 ms at hop 5 and stays there to the destination. A jump that persists marks where the delay is added, perhaps a long-distance link. A single high hop that drops back afterwards is usually just a router deprioritizing ICMP.

So: the office LAN is fine (hop 1 is fast), and the extra latency is in the provider path. You'd attach this trace and a pathping/mtr to the escalation.

How it breaks, and what to check

SymptomLikely causeWhat to check
Can reach some subnets, not otherswrong mask, missing route, or firewall between subnetsroute print / ip route; compare mask with a working machine
Works in office, not on VPNoverlapping home subnet, missing split-tunnel route, VPN firewall policyCheck home LAN range vs target; route print after connecting
Ping fails, website worksICMP blockedTest the actual port (module 5) instead of ping
Traceroute repeats the same two hopsrouting loopEscalate with trace output and time
Big pages/files hang, small ones fineMTU / blocked ICMP fragmentation-neededDF ping test with decreasing sizes (commands)

Commands

TaskWindowsmacOSLinux
Routing tableroute printnetstat -rnip route
Which route for one address?Find-NetRoute -RemoteIPAddress 10.8.3.4route get 10.8.3.4ip route get 10.8.3.4
Trace with per-hop losspathping 198.51.100.20mtr 198.51.100.20 (Homebrew)mtr -rwc 50 198.51.100.20
Find path MTU (1472 + 28 = 1500)ping -f -l 1472 192.0.2.10ping -D -s 1472 192.0.2.10ping -M do -s 1472 192.0.2.10
Force IPv4 / IPv6ping -4 / ping -6ping / ping6ping -4 / ping -6
At the help deskPing in order and stop at the first failure: own IP (stack OK) → gateway (LAN OK) → a public IP like 1.1.1.1 (routing/ISP OK) → a public name (DNS OK). Gateway works but public IP fails: the problem is past the LAN. Public IP works but names fail: go to DNS. Remember ping can be blocked; confirm with a port test before declaring something down.

Check yourself

Answer out loud first, then open the card.

How many usable hosts in a /27?

30 (32 addresses minus network and broadcast).

What network is 172.16.9.200/21 in?

Block size in the third octet is 8; 9 rounds down to 8 → 172.16.8.0/21 (172.16.8.0 – 172.16.15.255).

A packet matches 0.0.0.0/0 and 10.0.0.0/8. Which route is used?

10.0.0.0/8, the longer (more specific) prefix.

Why is a lone '* * *' hop in the middle of a traceroute usually harmless?

That router simply doesn't reply to probes; if later hops answer, it is still forwarding traffic.

5

Transport: TCP, UDP and why ports matter

MODULE 05 · ~6 min read

Getting data to the right application, either reliably or quickly, and what "refused" versus "timed out" is telling you.

IP gets a packet to the right machine. The transport layer gets it to the right program on that machine, using port numbers, and decides how careful to be about delivery. There are two main choices: TCP (reliable, ordered, connection-based) and UDP (lightweight, no guarantees).

In support, this layer is where "the server is up but the app won't connect" lives. Knowing a dozen common ports and the difference between a refused and a timed-out connection lets you point at the right team in one step.

Core ideas

Ports and the five-tuple

A port is a 16-bit number (0–65535). Servers listen on fixed, well-known ports; clients use a temporary ephemeral port picked by the OS (Windows uses 49152–65535 by default, Linux usually 32768–60999).

A connection is uniquely identified by five values: protocol, source IP, source port, destination IP, destination port. That's how one laptop can hold ten separate connections to the same server on 443: each has a different source port.

Ports worth knowing by heart

PortServicePortService
22 TCPSSH, SFTP443 TCP/UDPHTTPS (UDP = HTTP/3)
25 TCPSMTP server-to-server mail445 TCPSMB file sharing
53 UDP/TCPDNS587 TCPMail submission (client → server)
67/68 UDPDHCP636 TCPLDAPS
80 TCPHTTP993 TCPIMAP over TLS
88 TCP/UDPKerberos (AD logon)1433 TCPMicrosoft SQL Server
123 UDPNTP (time)3306 TCPMySQL
161 UDPSNMP (monitoring)3389 TCP/UDPRDP
389 TCP/UDPLDAP5060/5061SIP (VoIP signalling)

TCP: a reliable conversation

TCP opens with a three-way handshake: client sends SYN, server replies SYN-ACK, client sends ACK. Both sides now have starting sequence numbers.

  • Sequence numbers and acknowledgements: every byte is numbered; the receiver says "I have everything up to byte N". Missing data is retransmitted.
  • Flow control: the receiver advertises a window, how much it can accept right now, so a fast sender can't overwhelm a slow receiver.
  • Congestion control: TCP starts slowly and speeds up until it sees loss or delay, then backs off. That's why one lossy Wi-Fi link makes downloads crawl, not just "lose a bit".
  • Closing: each side sends FIN and the other acknowledges. A RST (reset) aborts a connection immediately.

Connection states you'll see in netstat: LISTEN, SYN_SENT (we asked, no answer yet), ESTABLISHED, TIME_WAIT (recently closed, normal), CLOSE_WAIT (the other side closed; our app hasn't. Lots of these means an app bug).

UDP: send and move on

UDP adds just ports, a length and a checksum to IP: no handshake, no acknowledgements, no retransmission, no ordering. If the application cares about loss, it handles it itself (DNS simply asks again after a timeout).

That's ideal when late data is useless: voice and video calls (RTP), online games, DNS lookups, DHCP, NTP, and QUIC, which rebuilds reliability on top of UDP for HTTP/3. A video call would rather skip 20 ms of audio than pause the whole conversation waiting for a resend.

Firewalls and NAT keep state

Stateful firewalls and NAT routers track connections. For TCP they see the handshake and the FIN. UDP has no "end", so the device keeps a mapping only for a timeout (often 30 s – a few minutes). When a VoIP phone stops sending keepalives and the mapping expires, incoming calls stop reaching it: one-way audio and "can call out but can't receive" tickets often trace back to this, or to SIP/NAT handling on the firewall.

Reading the error message

MessageWhat happened on the wireLook at
Connection refusedHost replied with RST: it's up, nothing listens on that portService stopped, wrong port, app bound to localhost only
Timed outNo reply at all to SYNFirewall dropping, wrong IP, host down, routing
Connection reset by peerRST in the middle of a sessionServer app crashed/restarted, a proxy or firewall cut it, idle timeout
No route to hostLocal routing or ICMP unreachableRoutes, VPN, gateway
Worked example

"I can't RDP to the reporting server"

PS> Test-NetConnection 10.40.2.15 -Port 3389
ComputerName           : 10.40.2.15
RemoteAddress          : 10.40.2.15
RemotePort             : 3389
PingSucceeded          : True
TcpTestSucceeded       : False

Ping works, so the network path is fine. The TCP test fails, so either Remote Desktop is off/stopped on the server, or a firewall (Windows Defender Firewall on the host, or a network firewall) blocks 3389. Next check: does the same test work from a colleague's machine in another subnet? If yes, it's a rule specific to the user's subnet or VPN pool. You've turned "RDP is broken" into "TCP 3389 blocked from 10.50.0.0/16 to 10.40.2.15", which a server or firewall team can act on.

How it breaks, and what to check

SymptomLikely causeWhat to check
"Connection refused"service stopped or on another portIs the service running? netstat/ss on the server
"Timed out" but ping worksfirewall dropping that portTest-NetConnection -Port / nc -vz from two different networks
App drops after N minutes idlefirewall/NAT/proxy idle timeoutNote exact interval; ask about session timeouts; enable keepalives
VoIP one-way audioNAT/firewall UDP mapping or SIP handlingTest on another network; escalate to voice team with call times
Large file copy slow on Wi-Fi onlyloss → TCP backs offCompare wired; check signal and retransmissions in capture

Commands

TaskWindowsmacOSLinux
Test one TCP portTest-NetConnection host -Port 443nc -vz host 443nc -vz host 443
Active connections + PIDnetstat -anonetstat -anv -p tcpss -tanp
Listening portsGet-NetTCPConnection -State Listenlsof -nP -iTCP -sTCP:LISTENss -tlnp / ss -ulnp
Quick HTTP port checkcurl.exe -v https://hostcurl -v https://hostcurl -v https://host
At the help deskTest the actual port, not just ping. Then read the error precisely: refused = the host answered, the service is the problem; timed out = something silently dropped it, look at firewalls and paths. Test from a second network or subnet to tell "blocked for everyone" from "blocked for this user's segment". UDP is hard to test with simple tools because silence is normal; rely on the application's own test (a DNS query, a test call).

Check yourself

Answer out loud first, then open the card.

What does TIME_WAIT mean, and is it a problem?

The connection closed recently; the OS holds the tuple briefly so late packets aren't confused with a new connection. Normal, unless there are tens of thousands of them.

Why does DNS normally use UDP?

Queries are tiny and single request/response; a handshake would double the latency. It falls back to TCP for large answers and zone transfers.

Refused vs timed out: which suggests a firewall?

Timed out (silent drop). Refused means the host actively replied that nothing is listening.

Mail client set to port 25 can't send from home. Why might 587 fix it?

Many ISPs block outbound 25 to stop spam; 587 is the authenticated submission port for clients.

6

DNS: turning names into addresses

MODULE 06 · ~6 min read

The distributed lookup system everything depends on, and the root cause of a large share of "the internet is broken" tickets.

People use names; networks route on IP addresses. DNS translates between them, and it's involved in nearly every connection: web, email, Teams, VPN, Active Directory logon, printers. When DNS breaks, it looks like everything broke, even though the network itself is fine.

The good news: DNS problems are very testable. With nslookup or dig and two different DNS servers you can usually prove within a minute whether DNS is at fault.

Core ideas

The namespace is a tree

Read a name from right to left: portal.corp.example.com. → the root (the invisible trailing dot) → the TLD com → the domain example.com → subdomains corp and portal. Each level can delegate the next level to other servers. Whoever runs example.com controls everything under it, without asking anyone.

Record types you'll meet

TypeHoldsSupport example
A / AAAAIPv4 / IPv6 addressportal → 203.0.113.10
CNAMEalias to another namewww → example.cdn-provider.net
MXmail servers, with prioritymail bouncing after a provider move
TXTfree text: SPF, DKIM, DMARC, domain verificationcompany mail landing in spam
NSwhich servers are authoritativedomain moved to a new DNS host
PTRreverse: IP → namemail servers checking senders
SRVservice location (host + port)AD clients find domain controllers via _ldap._tcp.dc._msdcs.corp.example.com
SOAzone details, serial, negative-cache timerarely touched directly

How a lookup actually happens

  1. The application asks the OS. The OS checks its cache and the hosts file (C:\Windows\System32\drivers\etc\hosts, /etc/hosts).
  2. If not found, the OS's stub resolver asks the DNS server it was configured with (from DHCP or VPN): a recursive resolver such as a domain controller, the router, the ISP, or a public service.
  3. The recursive resolver, if it has nothing cached, walks the tree: asks a root server who handles .com, asks a .com server who handles example.com, asks that authoritative server for the actual record.
  4. The answer comes back to the laptop, and every resolver along the way caches it for its TTL.

So "which DNS server am I using?" is always the first question: different resolvers can legitimately give different answers.

TTL, caching and "propagation"

Every record has a TTL in seconds. A resolver that cached portal → 203.0.113.10 with TTL 3600 will keep handing out that address for up to an hour, even if the record changed a minute later. That's what people call "DNS propagation": not a wave spreading, just caches expiring at different times. Admins lower the TTL a day before a planned migration to shorten this window.

Failures are cached too (negative caching). If someone looked up a name before it was created, that NXDOMAIN can stick for the zone's negative TTL.

Answers and what they mean

  • NOERROR with an address: DNS worked. If the site still fails, look elsewhere.
  • NXDOMAIN: the name doesn't exist (typo, record never created, wrong DNS suffix).
  • NOERROR, no answer: the name exists but not that record type (e.g. no AAAA).
  • SERVFAIL: the resolver couldn't get a valid answer: authoritative servers unreachable, or a DNSSEC validation failure.
  • REFUSED / timeout: the resolver won't serve you or isn't reachable (firewall, wrong IP, down).

DNS inside companies

  • Active Directory depends on DNS. Domain-joined PCs must use the company's DNS servers (usually domain controllers). Someone who sets 8.8.8.8 manually "to make the internet faster" breaks domain logon, Group Policy and file shares, because public DNS has never heard of corp.example.com.
  • Split DNS / split-horizon: the same name resolves to an internal IP inside (or on VPN) and a public IP outside, or internal names exist only on internal servers.
  • Search suffixes: typing just intranet works because the OS appends corp.example.com. Off VPN that suffix may be missing, so short names fail while full names work.
  • Encrypted DNS: DNS over HTTPS (port 443) and DNS over TLS (853). A browser using its own DoH provider can bypass company DNS, so internal names fail only in that browser.

DNS and email deliverability

Three TXT-based records decide whether a company's email is trusted: SPF lists which servers may send for the domain, DKIM publishes a key to verify message signatures, and DMARC tells receivers what to do when those checks fail. A newly added marketing tool or scanner that sends "as" the company without being in SPF/DKIM is a typical reason messages land in spam or bounce.

Worked example

Internal portal fails on VPN; public sites fine

C:\> nslookup portal.corp.example.com
Server:  dns.google
Address: 8.8.8.8
*** dns.google can't find portal.corp.example.com: Non-existent domain

C:\> nslookup portal.corp.example.com 10.20.0.2
Server:  dc01.corp.example.com
Address: 10.20.0.2
Name:    portal.corp.example.com
Address: 10.20.0.15

The record exists (the company DNS server answers). The laptop is simply asking the wrong server: 8.8.8.8 is set manually on the adapter, overriding what the VPN pushes. Set DNS back to automatic, reconnect, flush the cache, retest. Two lookups, one clear cause.

How it breaks, and what to check

SymptomLikely causeWhat to check
Names fail, IPs workDNS server unreachable or wrongnslookup name vs nslookup name 1.1.1.1
Old site appears after a migrationcached record within TTL; hosts file entryipconfig /displaydns; check hosts file; flush
Internal names fail on VPNVPN DNS not applied, manual DNS set, missing suffixipconfig /all DNS servers; try the FQDN
Domain logon/GPO errorsPC using public DNS instead of DCDNS servers must be DCs/company resolvers
Works in Edge, not in Firefox (or vice versa)browser using its own DoHBrowser DNS-over-HTTPS setting / enterprise policy
SERVFAIL for one domainDNSSEC failure or broken authoritative serversTry another resolver; dig +trace; report to domain owner
Company mail lands in spamSPF/DKIM/DMARC mismatchCheck message headers for spf= dkim= dmarc= results

Commands

TaskWindowsmacOSLinux
Basic lookupnslookup portal.example.comnslookup portal.example.comdig portal.example.com
Ask a specific servernslookup portal.example.com 1.1.1.1dig @1.1.1.1 portal.example.comdig @1.1.1.1 portal.example.com
Specific record typeResolve-DnsName example.com -Type MXdig example.com MX +shortdig example.com TXT +short
Follow the delegation—dig +trace example.comdig +trace example.com
Which resolvers am I using?ipconfig /all; Get-DnsClientServerAddressscutil --dnsresolvectl status
Show / clear cacheipconfig /displaydns; ipconfig /flushdnssudo dscacheutil -flushcache; sudo killall -HUP mDNSResponderresolvectl flush-caches
At the help deskProve or rule out DNS in two lookups: once against the configured server, once against a known-good one (company DC for internal names, 1.1.1.1 for public). Same answer from both and the site still fails? It isn't DNS. Different answers? Find out which is right and why the laptop is using the other. Check the hosts file for leftover entries from old projects.

Check yourself

Answer out loud first, then open the card.

A record was changed 10 minutes ago and a user still gets the old IP. Bug?

Usually not: their resolver or OS cached it and will keep it until the TTL expires. Flush local cache; the resolver cache may need time.

Why must domain-joined PCs use internal DNS servers?

AD clients find domain controllers via SRV records in the internal zone; public resolvers don't have them, so logon, GPO and Kerberos fail.

What does SERVFAIL suggest versus NXDOMAIN?

NXDOMAIN: the name doesn't exist. SERVFAIL: the resolver couldn't get a valid answer (authoritative servers down or DNSSEC failure).

Which record tells other servers where to deliver mail for a domain?

MX.

7

HTTP and HTTPS: how browsers and servers talk

MODULE 07 · ~6 min read

Requests, responses, status codes, cookies, caching and certificates: the layer users actually see.

Most business apps today are web apps, or desktop apps that talk HTTPS behind the scenes. HTTP is a simple text protocol of requests and responses; HTTPS is the same thing inside an encrypted TLS connection. When something goes wrong, the browser almost always tells you why, if you know where to look: the status code and the certificate.

This module gives you enough to read browser DevTools and curl output with confidence, and to tell "your laptop" problems from "their server" problems.

Core ideas

Anatomy of a URL

https://portal.example.com:443/reports/q3?region=emea#totals
└─┬─┘   └───────┬────────┘ └┬┘ └────┬────┘ └────┬────┘ └──┬──┘
scheme        host        port    path       query    fragment

The port is implied (443 for https, 80 for http) unless written. The fragment after # never leaves the browser. It's used by the page itself.

A real request and response

GET /reports/q3?region=emea HTTP/1.1
Host: portal.example.com
User-Agent: Mozilla/5.0 ...
Accept: text/html
Cookie: session=7f3a...

HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
Cache-Control: private, max-age=0
Set-Cookie: session=7f3a...; Secure; HttpOnly; SameSite=Lax
Strict-Transport-Security: max-age=31536000

<!doctype html> ...

A request is a method, a path, headers and sometimes a body. A response is a status code, headers and usually a body. The Host header lets one server/IP host many sites.

Methods and what they imply

  • GET read something; HEAD same but headers only (what curl -I sends).
  • POST submit/create (forms, logins, uploads).
  • PUT / PATCH replace / partially update; DELETE remove. Common in APIs.
  • OPTIONS used by browsers for CORS "preflight" checks before cross-site API calls.

Status codes: who is complaining?

CodeMeaningUsually points at
200 / 204OK / OK with no bodyWorking
301 / 308Moved permanentlyOld bookmark; redirect loops if misconfigured
302 / 307Temporary redirectNormal for SSO logins
304Not modified, use your cached copyCaching working
400Bad requestCorrupt cookie, oversized header, client bug
401Not authenticatedNot signed in / token expired
403Authenticated but not allowedPermissions, IP allowlist, proxy/WAF block
404Not foundWrong URL, removed page
407Proxy authentication requiredCorporate proxy (module 9)
429Too many requestsRate limiting
500Server errorApplication bug or crash
502 / 504Bad gateway / gateway timeoutSomething behind a proxy or load balancer (module 9)
503Service unavailableOverloaded, maintenance, no healthy backends

Rule of thumb: 4xx = the request (who you are, what you asked for), 5xx = the server side. A 5xx is rarely fixed on the user's laptop.

Cookies, sessions and SSO

HTTP has no memory between requests. After you log in, the server sets a cookie (or the app stores a token) that the browser sends with every later request. Single sign-on (SAML, OpenID Connect) bounces the browser between the app and the identity provider with redirects, setting cookies on each domain.

That's why "clear cookies for this site" or "try a private window" fixes so many login loops: a stale or corrupt session cookie, or a browser setting that blocks the cookies the SSO flow needs. Clearing all browsing data is rarely necessary; target the affected sites.

Caching

Browsers, CDNs and proxies cache responses according to headers like Cache-Control, ETag and Age. After a site update, a user might keep getting old scripts. A hard reload (Ctrl+Shift+R, Cmd+Shift+R on Mac) bypasses the browser cache; a CDN cache needs the site owner to purge it.

HTTPS: TLS and certificates

HTTPS wraps HTTP in TLS, which provides encryption (no one on the path can read it), integrity (no one can alter it) and authentication (the server proves who it is with a certificate).

A certificate binds a public key to names (listed in the Subject Alternative Name field) and is signed by a certificate authority. The browser follows the chain: server certificate → intermediate → a root CA that's already in the device's trust store. It also checks the validity dates against the local clock.

  • NET::ERR_CERT_DATE_INVALID: certificate expired, or the PC's clock/date is wrong.
  • ERR_CERT_COMMON_NAME_INVALID: you went to a name the certificate doesn't list (e.g. by IP, or an old alias).
  • ERR_CERT_AUTHORITY_INVALID: issuer not trusted: self-signed device page, missing intermediate, or a TLS-inspecting proxy whose root isn't installed.

Sites using HSTS don't allow clicking through the warning at all, which is deliberate.

HTTP versions

HTTP/1.1 handles one request at a time per connection, so browsers open several. HTTP/2 multiplexes many requests over one TLS connection. HTTP/3 runs over QUIC on UDP 443. Users don't choose; browsers negotiate. For support it matters mainly when a network blocks UDP 443 (HTTP/3 quietly falls back) or an old proxy mishandles HTTP/2.

Worked example

Login loop on the expenses app

  1. Open DevTools (F12) → Network, tick "Preserve log", reproduce.
  2. You see 302 to the identity provider, 302 back to the app, the app returns 302 to the identity provider again… forever.
  3. Try a private window: login works. So the service is fine and this browser profile has a bad cookie or a blocking setting/extension.
  4. Clear cookies for the app and identity-provider domains only; check that third-party cookie blocking or a privacy extension isn't stripping the SSO cookies. Retest.

If the private window also loops, it's not the browser. Note the time, user, and any correlation ID shown, and escalate to the app/identity team.

How it breaks, and what to check

SymptomLikely causeWhat to check
Certificate warning on every sitewrong system clock, or TLS inspection root missingCheck date/time and time sync; check issuer of the certificate
Certificate warning on one siteexpired cert, name mismatch, missing intermediateClick the padlock/view certificate; openssl s_client
Login loopstale cookie, blocked third-party cookies, clock skewPrivate window; clear cookies for those domains
403 for one user onlypermission/group missing, or IP/location policyCompare with a colleague; check group membership
Old version of page after releasebrowser or CDN cacheHard reload; ask owner to purge CDN
502 / 503 / 504server side, behind a proxy/LBStatus page; note time + error page branding; escalate

Commands

TaskWindowsmacOSLinux
Headers onlycurl.exe -I https://portal.example.comcurl -I https://portal.example.comcurl -I https://portal.example.com
Full conversation incl. TLScurl.exe -v https://portal.example.comcurl -v …curl -v …
Certificate dates, names, issuerBrowser padlock → certificateopenssl s_client -connect portal.example.com:443 -servername portal.example.com </dev/null | openssl x509 -noout -subject -issuer -datessame as macOS
Is the clock right?w32tm /query /statussntp time.apple.comtimedatectl
Fix time syncw32tm /resync (admin)System Settings → Date & Time → automaticsudo timedatectl set-ntp true
At the help deskGet the exact status code or certificate error, not "it doesn't work". DevTools → Network shows the failing request; a private window rules out cookies, cache and extensions in one step; another device or network separates "this laptop" from "this service". A wrong clock breaks HTTPS everywhere at once and is a two-minute fix.

Check yourself

Answer out loud first, then open the card.

401 vs 403?

401: you're not authenticated (log in). 403: you are, but you're not allowed.

A user's laptop is set to 2019. What breaks?

HTTPS everywhere (certificates look not-yet-valid or expired), Kerberos/AD logon, and MFA codes.

Why does a private window fix many web issues?

It starts with no cookies, no cache and usually no extensions, so stale sessions or extensions are ruled out.

What's checked when a browser validates a certificate?

Name matches (SAN), within validity dates, chains to a trusted root, not revoked (where checked), and the server proves it holds the key.

8

Beyond plain HTTP: TLS, WebSocket, MQTT, QUIC and friends

MODULE 08 · ~5 min read

What these names mean when they show up in a ticket, a log or a firewall request, plus the infrastructure protocols behind everyday office IT.

Plain request/response HTTP isn't enough for everything. Chat needs messages pushed instantly, sensors need tiny messages over bad links, video needs low latency. Several protocols fill those gaps, and they tend to appear in support work as "live updates don't arrive", "the device shows offline" or "can you open port 8883?"

The honest goal at help-desk level: recognize each one, know its port and transport, and know what kind of network setting typically breaks it. You don't need to implement them.

Core ideas

TLS beyond the browser

TLS secures far more than websites: email (IMAPS 993, SMTP with STARTTLS on 587), LDAPS (636), VPNs, MQTT over TLS, RDP. Versions: TLS 1.2 and 1.3 are current; 1.0 and 1.1 are deprecated and disabled by modern clients, which is why very old devices (printers, scanners, legacy apps) suddenly "can't connect" after updates.

Mutual TLS (client certificates) is where the client also presents a certificate. It's used for 802.1X Wi-Fi, some VPNs and device-trust checks. An expired machine certificate shows up as "the laptop can't join Wi-Fi/VPN, but my phone can".

WebSocket: a connection that stays open

A WebSocket starts as an HTTP request carrying Upgrade: websocket. The server answers 101 Switching Protocols and the same TCP connection becomes a two-way message channel (ws://, or wss:// over TLS). Chat apps, live dashboards, trading screens, collaborative editors and notification systems use it.

What breaks it: proxies or firewalls that don't support the upgrade, or that cut "idle" connections after a timeout. The page loads fine (that part is normal HTTP) but nothing updates live. Server-Sent Events (a long-lived one-way HTTP response) is a simpler cousin with similar failure modes.

MQTT: publish/subscribe for devices

MQTT is a small protocol for IoT: sensors, building systems, industrial equipment. Devices keep a TCP connection to a broker. A publisher sends a message to a topic like site1/lab/temp; the broker forwards it to every client subscribed to that topic. Publishers and subscribers never talk directly.

  • Ports: 1883 plain, 8883 TLS (sometimes WebSocket on 443).
  • QoS levels 0 (at most once), 1 (at least once), 2 (exactly once) trade reliability for overhead.
  • Retained messages give new subscribers the last known value; a last will message announces when a device disconnects unexpectedly.

"Device shows offline" often means its outbound 8883 is blocked on a new network, or its certificate expired.

QUIC and HTTP/3

QUIC runs over UDP 443 and builds in what TCP + TLS do separately: reliable streams, encryption and a faster handshake. One lost packet only stalls its own stream, and connections can survive a change of network (Wi-Fi to mobile). Browsers try HTTP/3 and fall back to TCP if UDP 443 is blocked, so it rarely causes outright failures; some firewalls deliberately block it so they can inspect traffic over TCP.

gRPC and APIs

Many internal services talk REST (JSON over HTTPS) or gRPC (binary messages over HTTP/2). They matter in support mostly through proxies: a proxy that downgrades to HTTP/1.1 breaks gRPC. Recognize the name, pass it to the right team.

Office infrastructure protocols

ProtocolPortWhat it does / typical ticket
SMBTCP 445Windows file shares, mapped drives. Blocked over the internet by most ISPs; needs VPN.
Kerberos88AD authentication. Fails with clock skew over ~5 minutes.
LDAP / LDAPS389 / 636Directory lookups; scanners and apps using "scan to email" or address books.
NTPUDP 123Time sync. Everything security-related depends on it.
RDPTCP/UDP 3389Remote desktop; UDP improves responsiveness.
SSHTCP 22Remote shell, secure file copy, tunnels.
SIP / RTP5060-5061 / UDP rangeVoIP call setup / the audio itself. One-way audio = RTP path.
SNMPUDP 161/162Monitoring switches, printers, UPS units.
IPP / raw printing631 / 9100Printing to network printers.
Worked example

Dashboard loads, but numbers never update

  1. DevTools → Network → filter WS. You see a wss://live.example.com/feed request with status (failed) or 400, instead of 101.
  2. On a phone hotspot it works: the app is fine; something on the company network interferes.
  3. The corporate proxy is the usual suspect: WebSocket upgrade not permitted, TLS inspection breaking it, or an idle timeout of 60 s cutting quiet connections.
  4. Escalate with the URL, time, the failed status, and the hotspot comparison. The fix is usually a proxy bypass or rule for that host.

How it breaks, and what to check

SymptomLikely causeWhat to check
Old printer/scanner can't send to cloud or mailTLS 1.0/1.1 no longer acceptedFirmware update; check supported TLS versions
Live updates never arriveWebSocket blocked/cut by proxy or firewallDevTools WS filter: look for 101; compare on hotspot
IoT device offline on new networkoutbound 8883/1883 blocked, cert expiredFirewall logs for device IP; device cert dates
Mapped drives fail from homeSMB 445 not reachable without VPNConnect VPN; confirm 445 test to file server
Logon errors after BIOS battery diedclock skew breaks Kerberos and TLSFix time, resync, retry

Commands

TaskWindowsmacOSLinux
Test TLS versions a server allows—openssl s_client -connect host:443 -tls1_2same; try -tls1_1 to see refusal
Test MQTT TLS portTest-NetConnection broker -Port 8883nc -vz broker 8883nc -vz broker 8883
Test SMB reachabilityTest-NetConnection files01 -Port 445nc -vz files01 445nc -vz files01 445
Kerberos ticketsklistklistklist
At the help deskWhen a feature is "half working" (page loads, live parts don't), think of a second protocol on a long-lived connection. Look for a 101 in DevTools, test on a different network, and check whether a proxy, TLS inspection or idle timeout sits in between. For devices, ask which port and protocol the vendor documents and test exactly that.

Check yourself

Answer out loud first, then open the card.

What HTTP status shows a successful WebSocket upgrade?

101 Switching Protocols.

In MQTT, does a sensor send data directly to the dashboard?

No. It publishes to a topic on the broker; the broker delivers to subscribers.

If a firewall blocks UDP 443, does HTTP/3 break websites?

Normally no; browsers fall back to HTTP/2 or 1.1 over TCP.

Why might a laptop's clock being 10 minutes off break domain logon?

Kerberos rejects tickets with more than about 5 minutes of clock skew by default.

9

Proxy servers: forward, reverse, PAC files and TLS inspection NEW

MODULE 09 · ~11 min read

The middlemen between users and the web. They explain a whole family of "works at home, not at the office" tickets, and most 502/504 errors.

A proxy is a server that accepts a connection from a client, then makes its own connection to the destination on the client's behalf. That's different from a router, which just forwards packets, and from NAT, which only rewrites addresses. A proxy actually terminates the conversation, so it can read (some of) it, make decisions, log it, cache it, or refuse it.

There are two families. A forward proxy sits in front of users and represents them to the internet: the corporate web proxy. A reverse proxy sits in front of servers and represents them to users: the thing behind most websites' HTTPS. Both appear constantly in support, often without anyone mentioning the word "proxy".

Core ideas

Forward proxy: speaking for the users

Companies route web traffic through a forward proxy (on-premises, or a cloud "secure web gateway") to:

  • filter by category or reputation (block malware and phishing sites, sometimes social media);
  • log who went where, for security investigations and compliance;
  • control egress: only the proxy is allowed out to the internet, so malware on a PC can't connect directly;
  • give the company one predictable exit IP, which partners and SaaS tenants can allowlist;
  • historically, cache common content (much less useful now that almost everything is HTTPS).

An explicit proxy is configured on the device (browser/OS settings, PAC file, or environment variables). A transparent or intercepting proxy is placed in the network path and grabs traffic without any client setting. The user may not know it exists at all.

How an explicit proxy handles HTTP and HTTPS

For plain http://, the browser sends the full URL to the proxy and the proxy fetches it:

GET http://intranet-old.example.com/news HTTP/1.1
Host: intranet-old.example.com

For https://, the browser asks the proxy to open a raw tunnel, then runs TLS through it end to end:

CONNECT portal.example.com:443 HTTP/1.1
Host: portal.example.com:443

HTTP/1.1 200 Connection established
(TLS handshake and encrypted HTTP now pass through the tunnel)

Without inspection, the proxy only learns the hostname and port (from CONNECT, plus the SNI field in the TLS hello) and how many bytes moved, not the pages or contents. Common proxy ports are 8080, 3128 and 80, but it's whatever the company chose.

Proxy authentication and 407

A proxy that needs to know who you are replies 407 Proxy Authentication Required. On domain-joined Windows machines this is usually invisible: the browser answers with the user's Windows logon via Kerberos or NTLM ("integrated authentication"). Cloud gateways often use an agent on the laptop or SSO instead.

When it isn't invisible, users see repeated username/password pop-ups, or command-line tools fail with "407". Typical causes: the device isn't domain-joined, the user's password changed and a cached credential is stale, or the tool simply doesn't support integrated auth.

How devices learn about the proxy: manual, PAC, WPAD

  • Manual: a host and port (proxy.corp.example.com:8080) plus a bypass list of addresses that should go direct.
  • PAC file: a small JavaScript file at a URL (e.g. http://pac.corp.example.com/proxy.pac) that the browser runs for every request to decide which proxy, if any, to use.
  • WPAD: automatic discovery of the PAC file via DHCP option 252 or the DNS name wpad.<your-domain>. Convenient, but if a network lets an attacker answer for wpad, they can redirect traffic, so many companies push an explicit PAC URL through policy instead.

A PAC file returns a string such as "PROXY proxy.corp.example.com:8080; DIRECT", a list tried in order. DIRECT means no proxy.

A PAC file, read line by line

function FindProxyForURL(url, host) {
  // 1. Plain host names (no dots) and internal domains go direct
  if (isPlainHostName(host) || dnsDomainIs(host, ".corp.example.com"))
    return "DIRECT";
  // 2. Private address ranges go direct
  if (isInNet(host, "10.0.0.0", "255.0.0.0") ||
      isInNet(host, "192.168.0.0", "255.255.0.0"))
    return "DIRECT";
  // 3. Real-time media should not be proxied
  if (shExpMatch(host, "*.media.example-meetings.com"))
    return "DIRECT";
  // 4. Everything else: proxy A, then proxy B, then give up
  return "PROXY proxy-a.corp.example.com:8080; PROXY proxy-b.corp.example.com:8080";
}

Two support lessons hide in there. First, if the PAC URL itself can't be downloaded (VPN down, DNS issue), browsers behave differently: some go direct, some fail. Second, a missing bypass rule sends an internal site to the proxy, which can't reach it, so the user sees a proxy error page for an internal URL. Note that isInNet() resolves names through DNS, so a slow DNS server makes every page load slow.

One machine, several proxy settings

It surprises people that "the proxy setting" isn't one setting:

  • Windows has the user-level setting (Settings → Network → Proxy, a.k.a. WinINET) used by browsers and most apps, and a separate WinHTTP setting used by system services and some agents. A browser that works while Windows Update, an EDR agent or a sync service fails is a classic WinHTTP gap. netsh winhttp show proxy reveals it.
  • macOS stores proxies per network service (Wi-Fi vs Ethernet can differ).
  • Command-line tools (curl, git, pip, npm, Docker) usually read HTTPS_PROXY, HTTP_PROXY and NO_PROXY environment variables, or their own config files, and ignore the OS setting.
  • Java, some Electron apps and older line-of-business apps have their own proxy fields.

Reverse proxy: speaking for the servers

A reverse proxy receives requests from the internet for www.example.com and passes them to backend servers that are never exposed directly. nginx, HAProxy, Apache httpd, IIS with ARR, cloud load balancers (AWS ALB, Azure Application Gateway) and CDNs such as Cloudflare all act as reverse proxies. They typically:

  • terminate TLS: the public certificate lives here;
  • route by host name or path (/api to the API servers, / to the web servers);
  • load balance and health-check backends;
  • cache, compress, rate-limit, and filter attacks (a WAF);
  • add headers such as X-Forwarded-For so the backend knows the real client IP, since otherwise every request seems to come from the proxy.

When a user hits an error page, its look often tells you which layer produced it: a CDN-branded page, a plain "nginx" page, or the application's own styled error each point to a different team.

502, 503, 504: errors made by the middleman

These codes are usually generated by the proxy, describing its conversation with the server behind it:

  • 502 Bad Gateway: the proxy couldn't get a valid response: backend refused the connection, crashed mid-response, returned garbage, or the TLS connection from proxy to backend failed.
  • 504 Gateway Timeout: the backend accepted the request but didn't answer before the proxy's timer ran out (slow query, overloaded app, firewall silently dropping between proxy and backend).
  • 503 Service Unavailable: the proxy or app deliberately refused: maintenance mode, overload protection, or no healthy backends left in the pool.

The user's laptop is almost never the cause. Useful facts for the escalation: exact URL, time with timezone, whether it's every request or intermittent, which error page/branding appeared, and any request ID shown on it.

TLS inspection (break and inspect)

Without inspection, a forward proxy can't see inside HTTPS. With TLS inspection it acts as a deliberate, authorized man-in-the-middle: it terminates the user's TLS session using a certificate it generates on the fly for portal.example.com, signed by the company's own inspection CA, then opens a second, genuine TLS session to the real site. It can now scan downloads and enforce data-loss rules.

This only works if the device trusts the company CA. Managed browsers do (the root is pushed by Group Policy or MDM). Things that keep their own trust store don't: Python (certifi), Java keystores, Node.js, git, some Electron apps, Docker. They fail with errors like unable to get local issuer certificate or PKIX path building failed. Apps that pin certificates refuse inspection entirely and need a bypass. Companies also commonly exempt categories such as banking and healthcare for privacy reasons.

How to spot it: view the certificate in the browser. If the issuer is something like "Corp Inspection CA" rather than a public CA, the traffic is being inspected.

Proxy, VPN, NAT, SOCKS: telling them apart

Works atSeesTypical use
NATIP/port rewriteheaders onlysharing one public IP
VPNwhole IP traffic in a tunneleverything routed into itremote access to the company network
HTTP proxyHTTP/HTTPS requestsURLs (HTTP), hostnames (HTTPS), contents if inspectingweb filtering, logging, egress control
SOCKS proxyany TCP (SOCKS5 also UDP)destination host/portgeneric relays, ssh -D tunnels
Reverse proxyin front of serversfull requests (it terminates TLS)TLS, routing, load balancing, WAF
Worked example

"Teams works, but Python says certificate verify failed"

A developer on the office network runs pip install requests and gets SSL: CERTIFICATE_VERIFY_FAILED … unable to get local issuer certificate. In the browser, PyPI loads fine.

  1. Check the browser certificate for pypi.org: issuer is "Corp Inspection CA". TLS inspection is in play.
  2. The browser trusts that CA via the OS store; Python uses its own bundle, which doesn't include it.
  3. Correct fix: point the tool at the company CA bundle (for pip, pip config set global.cert /path/to/corp-ca.pem; for many tools REQUESTS_CA_BUNDLE or SSL_CERT_FILE), following the company's documented process. Or request an inspection bypass for package repositories if policy allows.
  4. Wrong fix: disabling certificate verification (--trusted-host, verify=False, curl -k). It hides the symptom and removes the protection.

How it breaks, and what to check

SymptomLikely causeWhat to check
Site works at home/hotspot, not in officeproxy category block, missing bypass, or inspection breaking itRead the block page; check certificate issuer; try with proxy bypass if allowed
Repeated proxy login prompts407 with integrated auth failing: not domain-joined, stale credentialsCredential Manager; confirm domain trust; test another user
Browser fine, Windows Update/agent failsWinHTTP proxy not configurednetsh winhttp show proxy; import per company process
curl/git/pip fail, browser finetool ignores OS proxy, or doesn't trust inspection CAHTTPS_PROXY/NO_PROXY; company CA bundle
Internal site shows proxy error pageinternal domain missing from PAC bypass / NO_PROXYCheck PAC logic for that host
All browsing breaks on VPN connect/disconnectPAC URL unreachable from current networkDownload PAC URL manually; check DNS for its host
Choppy Teams/Zoom audio on office networkreal-time media forced through proxy/TCPVendor guidance is to send media direct (UDP); review PAC/bypass
502 / 504 on a company web appbackend down or slow behind reverse proxy/LBStatus page; note time, URL, error branding, request ID; escalate

Commands

TaskWindowsmacOSLinux
Show user proxy settingsSettings → Network & internet → Proxy; reg query "HKCU\Software\Microsoft\Windows\CurrentVersion\Internet Settings"scutil --proxyenv | grep -i proxy; gsettings get org.gnome.system.proxy mode
System/service proxynetsh winhttp show proxynetworksetup -getwebproxy Wi-Fi; networksetup -getsecurewebproxy Wi-Fi/etc/environment; service unit files
PAC URL in usesame registry key: AutoConfigURLnetworksetup -getautoproxyurl Wi-Figsettings get org.gnome.system.proxy autoconfig-url
Fetch through a proxy explicitlycurl.exe -x http://proxy.corp.example.com:8080 -I https://example.comsame with curlsame with curl
Bypass all proxies for a testcurl.exe --noproxy "*" -I https://example.comcurl --noproxy '*' -I …curl --noproxy '*' -I …
Who issued the cert I'm seeing?Browser padlock → certificate → issueropenssl s_client -connect example.com:443 -servername example.com </dev/null | openssl x509 -noout -issuersame
At the help deskWhen something works on a hotspot or at home but not in the office, think proxy before anything else. Collect: the exact error or block page, the certificate issuer, which app (browser vs CLI vs background service), and the proxy configuration actually in effect (netsh winhttp show proxy, scutil --proxy, environment variables). Never "fix" certificate errors by turning verification off; get the right CA or a sanctioned bypass.

Check yourself

Answer out loud first, then open the card.

Forward vs reverse proxy in one sentence each?

Forward: sits in front of clients and fetches from the internet for them. Reverse: sits in front of servers and receives internet requests for them.

Without TLS inspection, what does a proxy learn about an HTTPS visit?

The destination hostname and port (CONNECT/SNI), timing and byte counts, but not the URL path or content.

What's the difference between 502 and 504?

502: the proxy got an invalid response or no connection from upstream. 504: upstream didn't respond before the proxy's timeout.

The browser works but Windows Update fails behind the proxy. What do you check?

The WinHTTP proxy setting (netsh winhttp show proxy), which is separate from the browser/user setting.

Why do Python or Java tools fail on a network with TLS inspection?

They use their own trust stores, which don't contain the company's inspection CA.

10

Sockets: where applications meet the network

MODULE 10 · ~5 min read

What is listening, what is connecting, and why ports collide, at the level a support engineer needs.

Applications don't touch packets directly. They ask the operating system for a socket, an endpoint they can read from and write to, and the OS handles TCP, IP and the network card. Every browser tab, Teams call and OneDrive sync is a handful of sockets.

For support, sockets are where "is the service actually running and reachable?" gets answered. A few commands tell you which process owns a port, whether it's listening only on localhost, and whether connections are piling up.

Core ideas

What a socket is

A socket is identified by protocol, local IP and local port. A TCP connection is the pair of both ends (the five-tuple from module 5). The server's single listening socket on :443 spawns a new connected socket for each client, which is how one web server handles thousands of clients on the same port.

Server lifecycle: socket → bind → listen → accept

  1. socket(): ask the OS for an endpoint (TCP or UDP, IPv4 or IPv6).
  2. bind(): claim a local address and port, e.g. 0.0.0.0:8443. Fails with "address already in use" if someone else has it.
  3. listen(): start queuing incoming connection attempts.
  4. accept(): take the next waiting connection and get a new socket for that client.

The client side is shorter: socket() then connect() to the server's IP and port. The OS picks an ephemeral source port automatically.

Where it listens matters: 127.0.0.1 vs 0.0.0.0

  • 127.0.0.1:8080: only programs on the same machine can connect. Other PCs get "connection refused".
  • 0.0.0.0:8080 (or [::]:8080 for IPv6): all interfaces; reachable from the network if the firewall allows.
  • 10.40.2.15:8080: only via that one interface's address.

"It works on the server itself but not from my PC" is very often a service bound to localhost, or the host firewall (Windows Defender Firewall's "allow this app?" prompt answered "no").

Ports collide

Only one process can listen on a given IP + port + protocol. Two apps that both want 8080, or a crashed service that left a zombie process holding its port, produce "address already in use" / "port is already allocated". Find the owner (commands below), then decide which to stop or reconfigure.

TIME_WAIT, CLOSE_WAIT and running out of ports

After a TCP connection closes, the side that closed first keeps it in TIME_WAIT for a short period. That's normal. But a busy client opening and closing thousands of short connections per minute to the same server can run out of ephemeral ports ("port exhaustion"): new connections then fail intermittently, a pattern seen on middleware, proxies and badly written integrations.

A growing pile of CLOSE_WAIT on a server means the remote side closed but the application never closed its end: an application bug (leaked sockets), not a network issue.

Other kinds of socket

UDP sockets have no connection: the app just sends and receives datagrams on a port. Unix domain sockets (and Windows named pipes) let processes on one machine talk without the network, e.g. Docker's /var/run/docker.sock. TLS is a layer the application adds on top of a TCP socket.

Worked example

A five-line client and a ten-line server

# client: ask a web server for its headers
import socket
host = "example.com"
with socket.create_connection((host, 80), timeout=5) as s:
    s.sendall(f"HEAD / HTTP/1.1\r\nHost: {host}\r\nConnection: close\r\n\r\n".encode())
    print(s.recv(512).decode(errors="replace"))

# server: echo one line back to each client (run, then: nc 127.0.0.1 9000)
import socket
srv = socket.socket(socket.AF_INET, socket.SOCK_STREAM)   # socket()
srv.bind(("127.0.0.1", 9000))                              # bind(): localhost only!
srv.listen()                                               # listen()
while True:
    conn, addr = srv.accept()                              # accept()
    with conn:
        conn.sendall(b"you are " + str(addr).encode() + b"\n")

Try connecting to that server from another computer: refused, because it's bound to 127.0.0.1. Change it to 0.0.0.0 and allow the port in the firewall, and it works. That's the localhost-binding lesson in practice.

How it breaks, and what to check

SymptomLikely causeWhat to check
"Address already in use" on service startanother process owns the portFind PID by port; stop or reconfigure one
Works on server, refused from other PCsservice bound to 127.0.0.1Listening address in netstat/ss output
Works on server, timed out from other PCshost firewall blockingWindows Defender Firewall inbound rules; ufw/firewalld
Intermittent connection failures under loadephemeral port exhaustionCount TIME_WAIT; connection reuse/pooling in app
Server slowly stops respondingsocket leak (CLOSE_WAIT growing)Count CLOSE_WAIT over time; escalate to app owner

Commands

TaskWindowsmacOSLinux
Who owns port 8080?netstat -ano | findstr :8080 then tasklist /fi "pid eq 1234"lsof -nP -i :8080sudo ss -tlnp 'sport = :8080'
All listening socketsGet-NetTCPConnection -State Listen | sort LocalPortlsof -nP -iTCP -sTCP:LISTENss -tulnp
Count states(Get-NetTCPConnection -State TimeWait).Countnetstat -an | grep -c TIME_WAITss -s
Process for a PIDGet-Process -Id 1234ps -p 1234ps -p 1234 -o pid,cmd
At the help deskThree questions answer most "service unreachable" tickets: is a process listening on that port, on which address (localhost or all interfaces), and does the host firewall allow it? Check on the server first, then test from a client. Hand app owners precise facts ("port 8443 bound to 127.0.0.1 only") rather than "the network is blocking it".

Check yourself

Answer out loud first, then open the card.

A service shows LISTEN on 127.0.0.1:5000. Can a colleague's PC reach it?

No. It only accepts connections from the same machine; it needs to bind to 0.0.0.0 or the server's IP.

What does 'address already in use' mean?

Another socket already holds that IP/port/protocol; only one listener is allowed.

Many CLOSE_WAIT connections on a server. Network problem?

Usually not. The application isn't closing sockets after the peer closed: an app bug.

11

Troubleshooting toolkit and packet analysis

MODULE 11 · ~6 min read

The handful of commands that answer most network questions, a repeatable method, and how to read a packet capture.

Good network troubleshooting is less about knowing exotic tools and more about running simple ones in a sensible order and writing down what you saw. This module collects the toolkit, shows how to read the output, and introduces packet capture for when the commands aren't enough.

Core ideas

A method, not a guess

A structured approach (similar to the methodology CompTIA teaches) keeps you from jumping around:

  1. Identify the problem: exact symptom, who's affected, since when, what changed.
  2. Theorize: list likely causes, simplest and most common first.
  3. Test the theory with a command or swap; if wrong, go back to 2.
  4. Plan and fix, considering impact (don't reboot a switch at 10:00 on a Monday).
  5. Verify it works for the user, end to end, and prevent recurrence if you can.
  6. Document what you found and did.

ping, and what it really tells you

ping sends ICMP echo requests. Read three things: whether replies arrive, the time (latency) and its variation, and the loss percentage at the end. Run it for longer than four packets when diagnosing intermittent problems (ping -t on Windows, plain ping elsewhere; stop with Ctrl+C). The ttl= in replies is the remaining hop count, not the latency.

traceroute, pathping and mtr

Traceroute sends probes with TTL 1, 2, 3… and lists who replies "time exceeded". Windows tracert uses ICMP; macOS/Linux traceroute uses UDP by default (-I for ICMP, -T on Linux for TCP). Use -d/-n to skip slow reverse lookups.

Read it like module 4's example: silent hops are fine if later hops answer; a latency increase only matters if it persists to the end. pathping (Windows) and mtr combine traceroute with sustained loss statistics per hop, much better evidence for an ISP ticket.

The rest of the everyday toolkit

  • ipconfig / ifconfig / ip: addresses, gateway, DNS (module 2).
  • nslookup / dig / Resolve-DnsName: DNS (module 6).
  • arp -a / ip neigh: who's on the local segment (module 3).
  • netstat / ss / Get-NetTCPConnection: connections and listeners (module 10).
  • Test-NetConnection / nc: one TCP port (module 5).
  • curl: HTTP(S), headers, TLS, proxies, timings (modules 7, 9, 13). Built into Windows 10+ as curl.exe.
  • Built-in reports: netsh wlan show wlanreport (Windows Wi-Fi history), Wireless Diagnostics (macOS), Network Reset (Windows) as a last resort.

Packet capture: seeing the actual conversation

When commands disagree with the user's experience, capture the packets. Wireshark is the standard GUI; tcpdump works on macOS/Linux; Windows has the built-in pktmon. Capture on the affected machine, reproduce once, stop, then analyze.

Two kinds of filter: a capture filter limits what is recorded (BPF syntax: host 10.40.2.15 and port 443); a display filter hides what you don't want to look at afterwards (Wireshark syntax):

Display filterShows
dnsall DNS queries and answers
ip.addr == 10.40.2.15traffic to or from one host
tcp.port == 443one port
tcp.flags.syn == 1 && tcp.flags.ack == 0connection attempts only
tcp.flags.reset == 1resets (refusals, aborted sessions)
tcp.analysis.retransmissionlost/resent data
tls.handshake.type == 1TLS Client Hello (shows SNI hostname)

Reading common patterns

  • SYN, SYN, SYN, nothing back: something drops the traffic (firewall, wrong IP). Matches "timed out".
  • SYN then RST from server: port closed. Matches "refused".
  • Handshake, Client Hello, then RST or alert: TLS problem (version, inspection, certificate).
  • Many retransmissions and duplicate ACKs: loss on the path, often Wi-Fi.
  • DNS query with no response, then retry to a second server: slow or dead first resolver, a hidden cause of "slow internet".

Right-click a packet → Follow → TCP Stream to see one conversation on its own.

Ethics and privacy

Only capture on devices and networks you're authorized to inspect, ideally with the user aware. Captures can contain passwords on legacy plain-text protocols, session cookies and personal data. Store them as sensitive and delete them when the ticket closes.

Worked example

Writing up what you found

Summary:  Users on floor 3 (VLAN 30) cannot open https://hr.corp.example.com
Since:    2026-10-05 09:10 PT, after network change CHG-1234 (per change calendar)
Scope:    floor 3 only; floor 2 and VPN users fine
Tests:    ping 10.30.0.1 (gateway) OK, 0% loss
          nslookup hr.corp.example.com -> 10.20.0.40 (correct)
          Test-NetConnection 10.20.0.40 -Port 443 -> TcpTestSucceeded False
          capture: SYN retransmitted 3x, no SYN-ACK
Theory:   firewall rule blocks 10.30.0.0/24 -> 10.20.0.40:443
Ask:      network team to review rules changed in CHG-1234

That ticket can be acted on without anyone calling the user back. That is the real deliverable of troubleshooting.

How it breaks, and what to check

SymptomLikely causeWhat to check
Ping fine, users say "slow"loss or jitter not visible in 4 pings; slow DNS; app slowLong ping run; DNS timing; curl timing breakdown
Traceroute stops at a hopfirewall drops probes beyond, or real outageTry TCP traceroute; test the actual service port
Intermittent failures you can't reproducetiming-dependent: Wi-Fi roaming, DHCP renew, IdP outageAsk user for exact times; leave a capture or continuous ping running
Capture shows nothing for the targetcapturing on wrong interface, or traffic goes via VPN/proxyPick the right adapter (VPN virtual adapter); check proxy

Commands

TaskWindowsmacOSLinux
Continuous pingping -t 10.30.0.1ping 10.30.0.1ping 10.30.0.1
Trace without DNS lookupstracert -d hosttraceroute -n hosttraceroute -n host (-T -p 443 for TCP)
Capture to a filepktmon start --capture --pkt-size 0 -f cap.etl, pktmon stop, pktmon etl2pcap cap.etl -o cap.pcapngsudo tcpdump -i en0 -w cap.pcap host 10.40.2.15sudo tcpdump -i eth0 -w cap.pcap port 53
Quick read of DNS trafficWireshark filter dnssudo tcpdump -i en0 -n port 53sudo tcpdump -i any -n port 53
At the help deskWrite down the exact command and its output, with the time, every time. A capture or a pathping attached to an escalation saves a round of "can you reproduce it?". A single slow ping reply is normal; a pattern of loss or spikes over a few minutes is evidence.

Check yourself

Answer out loud first, then open the card.

Capture filter vs display filter?

Capture filter limits what gets recorded (BPF syntax); display filter only hides packets in an existing capture (Wireshark syntax).

A capture shows SYN sent three times and no reply. What does the user see?

A timeout. Something is dropping the connection attempt (firewall, wrong IP, host down).

Why prefer pathping/mtr over a single traceroute for an ISP ticket?

They measure loss and latency per hop over many samples, so they show a sustained problem instead of one snapshot.

12

Network performance: why it feels slow

MODULE 12 · ~6 min read

Bandwidth, latency, jitter and loss are different problems with different fixes. Measure which one you have.

"Slow" is the vaguest ticket there is. It might mean a download that takes too long (bandwidth), a remote desktop that lags behind the mouse (latency), a video call that breaks up (jitter and loss), or a page that sits blank for five seconds before loading quickly (DNS or the server). The fix for one does nothing for another, so the job is to measure which one it is.

Core ideas

Five numbers, five meanings

MetricUnitPlain meaningHurts most
BandwidthMbpssize of the pipebig downloads, backups, uploads
ThroughputMbpswhat you actually get through itsame
Latency (RTT)mstime for a round tripRDP/VDI, page loads, chatty apps
Jittermshow much latency variesvoice and video
Packet loss%data that never arriveseverything; TCP slows down sharply

Rough targets for good voice/video, commonly cited from ITU guidance and vendor docs: one-way latency under ~150 ms, jitter under ~30 ms, loss under ~1%.

Where latency comes from

Light in fiber travels about 200 km per millisecond. Los Angeles to London is roughly 8,800 km in a straight line, so even a perfect path costs ~45 ms each way, ~90 ms round trip, and real routes are longer. No bandwidth upgrade changes physics. On top of distance come queuing (waiting in busy router buffers), processing (firewalls, inspection proxies, VPN encryption) and the last link (Wi-Fi retries can add tens of milliseconds).

A VPN that sends a user in Istanbul through a concentrator in California before reaching a server in Frankfurt adds a huge detour. Split tunnelling or a nearer VPN gateway fixes that, not more bandwidth.

Why loss is worse than it sounds

TCP treats loss as a sign of congestion and slows down. A link with just 1–2% loss can cut a single download's throughput dramatically, and the further away the server, the worse the effect, because recovery takes round trips. That's why "fast speed test, but file copies to the other office crawl" is a real combination: the speed test uses many parallel connections to a nearby server.

Wi-Fi: the usual suspect

  • Signal and distance: weaker signal → lower data rate → more airtime per packet.
  • Shared airtime: Wi-Fi is a shared medium; one slow, distant client uses more air and slows everyone on that AP.
  • Interference and channel overlap: neighboring networks, microwaves, Bluetooth on 2.4 GHz.
  • Sticky clients: a laptop that stays associated to a far AP instead of roaming.

The single most useful test: run the same measurement wired. If wired is fine, it's the Wi-Fi.

Congestion, bufferbloat and QoS

When a link is full, routers queue packets. Oversized queues (bufferbloat) make latency soar whenever someone uploads a big file: calls stutter "every time Bob's backup runs". Test it with a latency measurement during a speed test (many speed tests now report "loaded latency"; macOS has networkQuality).

QoS marks and prioritizes traffic (voice is commonly marked DSCP EF / 46) so calls go first when the link is busy. It only helps where the network is configured to honor the markings, typically inside a company WAN, not across the public internet.

MTU and fragmentation

Ethernet's MTU is 1500 bytes. VPNs and some tunnels add headers, shrinking what fits. If a packet is too big and the "don't fragment" flag is set, a router should reply with ICMP "fragmentation needed". When firewalls block that, connections hang on large transfers while small requests work: the login page loads, the report download stalls. Test with the DF ping from module 4; fixes are on the network side (MSS clamping, allowing ICMP) or lowering the VPN adapter MTU.

Not the network at all

Plenty of "slow network" tickets are the endpoint or the server: a laptop pegged at 100% CPU by an update or scan, a full disk, an overloaded SaaS tenant, a slow DNS resolver adding a second to every new hostname, or a browser with dozens of extensions. A curl timing breakdown (module 13) shows whether time is spent in DNS, connecting, TLS, or waiting for the server.

Worked example

"Teams calls are choppy every afternoon"

  1. Pin down when: the user says roughly 14:00–16:00, on Wi-Fi in a meeting room.
  2. Measure during the problem: a continuous ping to the gateway shows 3 ms most of the time with spikes to 180 ms and 2% loss; the same test wired in that room shows a clean 1 ms.
  3. Look at the Wi-Fi: signal is −72 dBm on 2.4 GHz; the room fills with 15 people at that time.
  4. Outcome: move the user to 5 GHz or wired for calls now; report coverage/capacity for that room to the network team with the numbers and times.

Note what you didn't do: upgrade the internet plan. The internet link was never the bottleneck.

How it breaks, and what to check

SymptomLikely causeWhat to check
Speed test fine, calls choppyjitter/loss, often Wi-FiLong ping during calls; wired comparison
Fine in morning, slow at 14:00congestion: shared uplink, backups, busy APRepeat measurements at both times; check uplink graphs
Slow only on VPNdetour via distant gateway; MTU; inspectionTrace with and without VPN; DF ping; split-tunnel policy
Small pages load, big downloads hangMTU/PMTUD blackholeDF ping test; lower VPN MTU; escalate MSS clamping
First visit to each site slow, then fineslow DNS resolvernslookup timing against two resolvers
One desk slow, neighbors finecable at 100 Mbps/half duplex, NIC driverLink speed; swap cable; update driver

Commands

TaskWindowsmacOSLinux
Latency + loss over timeping -n 200 10.30.0.1ping -c 200 10.30.0.1ping -c 200 10.30.0.1
Per-hop losspathping -n hostmtr -n hostmtr -rwn -c 100 host
Throughput between two machinesiperf3 -c 10.30.0.50iperf3 -c 10.30.0.50iperf3 -s on the other end
Loaded latency / responsivenessspeed test that reports loaded latencynetworkQuality -v—
Interface errors/dropsGet-NetAdapterStatisticsnetstat -I en0 -dip -s link show eth0
Wi-Fi rate and signalnetsh wlan show interfacesOption-click Wi-Fi iconiw dev wlan0 link
At the help deskMeasure, don't guess. Compare wired vs Wi-Fi on the same machine, run a ping for a few minutes to see jitter and loss, use pathping/mtr to see where it starts, and repeat at a different time of day. Put numbers and timestamps in the ticket. "Wi-Fi at −72 dBm on 2.4 GHz, 2% loss at 14:10, wired 0%" is actionable; "Wi-Fi slow" isn't.

Check yourself

Answer out loud first, then open the card.

Would doubling bandwidth fix a laggy remote desktop to another continent?

Probably not. RDP lag is mostly latency (distance, detours, Wi-Fi), which bandwidth doesn't reduce.

What is jitter and why does voice care?

Variation in packet delay. Voice plays audio in real time, so late or uneven packets cause gaps and robotic sound.

Small web pages load over VPN but large downloads hang. Likely cause?

An MTU/Path MTU Discovery problem, often because ICMP 'fragmentation needed' is blocked.

13

Networking behind large services

MODULE 13 · ~5 min read

Load balancers, CDNs, regions and dependencies: why big platforms fail partially, and how to escalate usefully.

Microsoft 365, Google Workspace, Salesforce or your company's own web app don't run on "a server". They run on many servers in many places, behind layers of DNS routing, CDNs and load balancers, often depending on other services such as an identity provider. That design makes them very reliable overall, but it also makes failures partial and confusing: one region, one feature or one ISP path can break while everything else works.

For support, the skill is recognizing a partial outage, proving whether it's "us" or "them", and handing the vendor or internal team the evidence they need.

Core ideas

Load balancers

A load balancer spreads requests across a pool of identical servers and continuously health-checks them, removing any that fail. Layer 4 balancers work on TCP/UDP connections; layer 7 balancers understand HTTP and can route by host or path (they're reverse proxies, module 9). Common algorithms: round robin, least connections, or hashing on the client.

Sticky sessions keep a user on the same server. When a sticky server is pulled out, those users may be logged out or lose a half-filled form, which explains "it kicked me out, then worked when I logged back in".

CDNs and caching at the edge

A content delivery network keeps copies of static content (images, scripts, video, software updates) at edge locations near users and forwards the rest to the origin. Benefits: speed, less load on the origin, and absorbing attacks. Side effects you'll see: users in different cities getting different versions until caches expire, and a CDN-branded error page when the origin behind it is down.

Response headers tell you what happened: Age, and vendor headers such as x-cache: HIT or cf-cache-status: HIT.

Getting users to the nearest place: GeoDNS and anycast

GeoDNS answers the same name with different IPs depending on where the request seems to come from (usually the resolver's location). Anycast announces the same IP address from many locations, and internet routing delivers you to the closest one. That's how 1.1.1.1 or a CDN IP is "everywhere". Consequence: a user whose DNS resolver is far away (e.g. corporate DNS in another country, or a VPN) may be sent to a distant edge and get slower service than colleagues.

Regions, zones and replication

Cloud providers group data centers into regions (geographic areas) and availability zones (separate facilities within a region). Services replicate data between them. Replication across distance takes time, so some systems are eventually consistent: a change saved in one place appears elsewhere a few seconds later. "I added the user but they can't see the shared folder yet" is often this, or a scheduled sync (e.g. directory sync running every 30 minutes).

Dependencies and cascading failures

Modern apps call many other services. If the identity provider (Entra ID, Okta, Google) has an outage, every SSO app fails at login while sessions already open keep working: the "nobody can log in, but people already in are fine" pattern. Internal services use timeouts, retries and circuit breakers; when retries pile up, a small slowdown can grow into a wide outage. Rate limiting (429) protects services from that.

Is it us or them?

  1. Status pages: the vendor's official status page and the admin-center service health. Not always instant, but check first.
  2. Change the network: try a phone hotspot. Works there → our network/proxy/ISP path; fails there too → likely the service or the account.
  3. Change the user/device: another account on the same machine, same account on another machine.
  4. Compare locations: other offices, VPN vs not, different DNS resolver (which may route to a different edge).

What to put in an escalation

Vendors and platform teams can't search "it's broken". They need: exact error text or screenshot; time with timezone; user, tenant or account; location and network (office, ISP, VPN); what you already ruled out; and any request / correlation / trace ID shown on the error page or in DevTools response headers. Those IDs let an engineer find the exact failed request in their logs.

Worked example

Where is the time going?

$ curl -o /dev/null -s -w "dns %{time_namelookup}s  connect %{time_connect}s  tls %{time_appconnect}s  first-byte %{time_starttransfer}s  total %{time_total}s\n" https://app.example.com/
dns 0.004s  connect 0.021s  tls 0.058s  first-byte 3.912s  total 3.950s

DNS, TCP and TLS together take under 60 ms; the server then takes almost four seconds to send the first byte. The network path is healthy and the delay is in the application or its backends. That's evidence for the app owner, not the network team. (On Windows use curl.exe and -o NUL.)

How it breaks, and what to check

SymptomLikely causeWhat to check
Works for some users, not othersdifferent regions/edges, sticky server, account-specificCompare location, network, account; note who works
Nobody can log in to SaaS apps; open sessions OKidentity provider outage or federation issueIdP status page; admin center health
Changes not visible for others yetreplication or sync delay, CDN cacheWait one sync interval; check sync logs; hard reload
Slow only from one officeISP path, far DNS resolver → distant edgetraceroute from there; compare resolvers; ISP ticket
CDN error pageorigin down or unreachable from CDNVendor status; note ray/request ID on page

Commands

TaskWindowsmacOSLinux
Which edge/IP do I get?nslookup app.example.comdig +short app.example.comdig +short app.example.com
Compare through another resolvernslookup app.example.com 1.1.1.1dig @1.1.1.1 +short app.example.comsame
Cache and server headerscurl.exe -sI https://app.example.comcurl -sI https://app.example.comsame
Timing breakdowncurl.exe -o NUL -s -w "%{time_starttransfer}" https://app.example.comsee examplesee example
At the help deskStatus page first, then change one thing at a time: network (hotspot), device, account, location. When escalating, include exact error text, time with timezone, location/network, what you ruled out, and any request or correlation ID. "Works on hotspot, fails on office network, request ID 8f2c…" turns a vague complaint into a precise problem.

Check yourself

Answer out loud first, then open the card.

Users can't log in to three different SaaS apps at once. Where do you look first?

The shared dependency: the identity provider (SSO) status and sign-in logs.

Why might two colleagues in different cities see different versions of a website?

They're served by different CDN edges whose caches expire at different times.

A curl timing shows tiny DNS/connect/TLS times but a 4-second first byte. Network or app?

Application/backend. The network part finished quickly; the server was slow to respond.