How the web works, DevTools, and your course repo
Agenda Slides Open fullscreen β
Open slides ↗Post-Class Notes
TL;DR#
Today we followed one page from the address bar to the screen: DNS, request, response, status code, and render. Then we opened DevTools on a real site and counted the files it actually downloads. In the last stretch we created the oim3690 course repository from your own computer and pushed it up, which is the reverse of Tuesday. When a number surprises you, measure it yourself before you accept an explanation for it.
What we did today#
- Three review polls on Tuesday's agent discussion, then a guess about how many files
www.babson.eduloads. - DevTools: Elements, Network, and the selector tool, on a real site.
- What happens between pressing Enter and seeing a page, one step at a time.
- Created the
oim3690course repository, published it, wrote anindex.html, and pushed.
A page is many files#
The poll asked how many separate files your browser downloads to open the Babson homepage. Most of the room said more than 100, and that turned out to be right, but the interesting part was that the first measurement disagreed.
With an ad blocker running, the Network tab showed about 90 requests. With the blocker turned off and the page reloaded, it went to nearly 200. Roughly half of a large commercial page is advertising and tracking that a blocker never fetches at all. This is also why the counts around the room did not match: your extensions, your cache, and your cookies all change what gets downloaded.
Read the number with that in mind:
- The same URL does not produce the same page on two different machines.
- A count in DevTools is a measurement of your browser on your network, so read it that way.
DevTools#
Open it from the Chrome menu under More Tools, or with Ctrl+Shift+I (Windows) or Cmd+Option+I (macOS). You will use it to inspect everything for the rest of this course, so it is worth 20 minutes of poking around on your own.
Elements shows the page as a tree of tags. Hover the selector button (top left of the panel), move the mouse over the page, and the matching code highlights. You can edit any text there and watch the page change. That change lives only in your browser, and it disappears when you reload, because nothing was sent back to the server.
Network shows every file the browser asked for. Tick Disable cache before reloading, otherwise some rows were never actually downloaded and the count is not honest. The first row is the HTML document itself, and its response is the source code of the page. Everything after it is CSS, JavaScript, images, and fonts.
Console is where JavaScript reports itself. We will use it properly later in the semester. There is something hidden in the console on the course site, if you want to go looking.
Right-click and View page source does something slightly different from Elements. Source is the raw text the server sent. Elements is the live tree your browser built from that text and has been changing ever since. On a simple page they look the same. On a complicated one they do not.
This is also how you learn from sites you like. Open one, look at how it is built, and ask your AI what a tag you have never seen does. Borrowing a technique is how everyone learns this. Copying a whole page is a different thing, so put your own ideas into what you take.
What every page has in common#
Open View page source on any two sites and the same shapes appear. Every line is wrapped in angle brackets. The first tag is <html> and the last is </html>. Just inside comes <head>, and after it closes, <body> opens and runs to the end.
<head>holds information about the page: the title, links to CSS files, and script tags pointing at JavaScript.<body>holds everything a visitor actually sees.
That is the whole structure of an HTML document, and the browser turns it into the tree you saw in Elements.
What happens when you press Enter#
The example on the slides was info.cern.ch, the first website ever put on the internet, which is still online. A URL has parts worth naming.
- Protocol:
https://. The rules for how the page is requested and returned. Thesis a security layer, and Chrome warns you when a site does not have it. - Domain:
www.babson.edu. The name of the machine you are asking. The last piece (.edu,.com,.io) is the top-level domain, and it is sold by registrars. Domains are not case sensitive and they cannot contain spaces, which is why a space inside a URL shows up as%20. - Path: everything after the domain. When you leave it out, the server hands you
index.html. That is exactly why your personal site needs a file with that name.
Then, in order:
- Find the address. The domain is a name, and the network needs a number.
nslookup www.babson.eduin a terminal returns the IP addresses behind it, usually several, because a big site runs on many machines. The system that keeps this lookup table is DNS, the Domain Name System. - Travel there.
tracert(Windows) ortraceroute(macOS) lists every machine between you and the destination. Google took only a handful of hops through Stamford and New York, which tells you Google has data centers near Boston, so nothing has to cross the country. The first address in everyone's list is the router they are sitting next to, which is why the lists in the room did not match. - Ask for the page. Your browser sends a request. The server sends back a response, and the first thing in it is the HTML.
- Read the status code.
200means the request worked.404means that page does not exist.301means moved permanently, which is a redirect. We triedbabsonrejects.comand it returns a 301 that sends you to Bentley. - Draw the page. The browser parses the HTML into a tree called the DOM (Document Object Model), does the same for the CSS, and paints the result. A modern browser is one of the most complicated pieces of software ever written, and this is most of what it does.
curl is worth knowing about even if you never use it yourself. It fetches a page from the command line with no browser involved: curl -I <url> shows the status code, and curl <url> prints the whole response. This is how an AI agent reads the web. An agent does not open Chrome, so anything it does with a website goes through a tool that works like this one.
Static sites and sites with a backend#
Some sites store information about you between visits. Instagram knows who you are, keeps your posts, and shows a different page to every visitor. That requires a database behind the site, which is what people mean by a backend.
Your site does not have one, and for the first half of this semester it will not need one. HTML, CSS and JavaScript build the front end, which is the part that runs in the visitor's browser. Backend frameworks come later in the semester if there is room.
Your course repository#
Tuesday went from GitHub down to your laptop. Today went the other way: the folder is made on your own machine first, then published. Every step, from New repository through commit and push, is on today's slides.
The automated checks fail later if either of these steps goes wrong. The repository has to be named oim3690 in lowercase, and "Keep this code private" has to be unchecked when you publish it.
The dropdown at the top of GitHub Desktop switches between repositories. You now have two, your personal site and oim3690, so check which one is selected before you commit.
Weekly notes#
Create a folder called logs inside oim3690, and a file inside it called wk01.md. One file per week, all semester. Write a few lines about this week:
- What you worked on.
- What did not work.
- Something new you picked up, in class or anywhere else.
- Something you noticed about working with AI this week.
A blank line is better than an invented one. These files are part of every checkpoint, and in December you will have 14 of them describing a semester you will not otherwise remember in detail. They are written in Markdown, which is the formatting used by every .md file you have opened so far, and we will cover the syntax properly next class.
Useful links#
- Getting Started and GitHub Pages, if any part of your setup is still incomplete.
- Weekly notes, for what belongs in
wk01.mdwhen a prompt comes up blank. - Live Server on the VS Code marketplace.
- info.cern.ch, the first website, still served from its original address.
- Chrome DevTools documentation, if you want to go further than we did.
Before Tuesday#
- Make sure
oim3690is public, and hasindex.htmlpushed to it. If you did not get to the file in class, ask your AI to generate a simple page, paste it intoindex.html, save, and push. - Write
oim3690/logs/wk01.md, with the four prompts above answered from this week. Both halves count in Checkpoint 1: the file being there, and it saying what actually happened. - Submit your
oim3690repository URL on Canvas. Due Monday 9/07, earlier than the rest of this list. It is how the automated checks find your work. - Install the Live Server extension if you have not yet.
- Keep building your personal site. It stays yours all semester.
Add/drop closes Friday 9/04.
Checkpoint 1 is due Saturday 9/12. It is checked directly on GitHub and there is nothing to hand in. From these first two weeks it looks for your personal site live at <your-github-username>.github.io with an index.html, your oim3690 repository public with a README.md and an index.html, and oim3690/logs/wk01.md. Next week adds to that list, and the full version goes on Canvas.
Next class we start writing HTML properly, and Markdown gets the time it did not get today. If something is broken, send me a message before Tuesday or find me at the start of class.