Operating Systems 2026F: Tutorial 2: Difference between revisions

From Soma-notes
No edit summary
 
(16 intermediate revisions by the same user not shown)
Line 1: Line 1:
'''This tutorial is still in development.'''
In this tutorial we're going to look at how processes work at a low level: how they make system calls & library calls, how C and assembly compare, and and how memory is laid out.
 
In this tutorial you will be learning more about how processes work on Linux.


==Getting Started==
==Getting Started==


For this tutorial, you need to get access to a Linux or UNIX machine.  We strongly suggest you use an SCS Openstack instance (see below).  You'll need access to a system for the entire semester, ideally the same one.
For this tutorial, you need to get access to a Linux or UNIX machine, and LinuxOnTab isn't enough.  We strongly suggest you use an SCS Openstack instance (see below).  You'll need access to a system for the entire semester, ideally the same one.


The concepts covered below are mostly part of standard UNIX/Linux tutorials.  Feel free to consult one or more of them.  However, remember that you are trying to build a conceptual model of how things work.  Thus, don't memorize commands; instead, try to understand how things fit together, and ask questions when things don't work as expected!
The concepts covered below are mostly part of standard UNIX/Linux tutorials.  Feel free to consult one or more of them.  However, remember that you are trying to build a conceptual model of how things work.  Thus, don't memorize commands; instead, try to understand how things fit together, and ask questions when things don't work as expected!
Line 12: Line 10:


Feel free to discuss this tutorial on Teams in the Tutorials channel.
Feel free to discuss this tutorial on Teams in the Tutorials channel.
===Openstack===


'''Again, for emphasis: don't take snapshots!''' (see below)
'''Again, for emphasis: don't take snapshots!''' (see below)
Line 23: Line 19:
Create a VM on the new SCS openstack cluster at [https://openstack-stein.scs.carleton.ca openstack-stein.scs.carleton.ca] and do your work there.  While you don't need a persistent VM for this lab, it will be important for future tuturials - and it is nice to have your work stick around when you leave the lab.
Create a VM on the new SCS openstack cluster at [https://openstack-stein.scs.carleton.ca openstack-stein.scs.carleton.ca] and do your work there.  While you don't need a persistent VM for this lab, it will be important for future tuturials - and it is nice to have your work stick around when you leave the lab.


'''To access Openstack''' you must be on the Carleton network, so make sure to VPN in!
'''To access Openstack''' you must be on the Carleton network, so make sure to [https://carleton.ca/its/vpn-access-to-campus/ VPN in] if you aren't on the Carleton computer or WiFi network.


'''To get added to the COMP 3000 project''', you need to [http://www.scs.carleton.ca/webacct change your SCS password] (run newacct) in order to update your account to have the right entitlements.
'''To get added to the COMP 3000 project''', you need to [http://www.scs.carleton.ca/webacct change your SCS password] (run newacct) in order to update your account to have the right entitlements.
Line 29: Line 25:
===Setting up and connecting a VM===
===Setting up and connecting a VM===


Create a VM on the SCS openstack cluster as shown in lecture. Make sure do the following:
Create a VM on the SCS openstack cluster as shown in the SCS documentation. Make sure do the following:
* Choose the comp3000-2026f-* snapshot image ONLY (others won't have the right software for later tutorials), use the latest one.
* Choose the COMP3000-F26.2026-* instance snapshot image ONLY (others won't have the right software for later tutorials), use the latest one.
* Add the ping-ssh-egress security group, and
* Add the ping-ssh-egress security group, and
* Associating a floating IP address.
* Associating a floating IP address.
Line 55: Line 51:


==Background==
==Background==
A process on UNIX-like systems are separated from each other: they run in their own address space - pointers can only refer to code and data in that program, not in other programs.  Because programs are separated, they can't access anything external to them without help.
Here we're going to talk about libraries and system calls, two ways programs gain access to additional functionality.  Libraries are for external code that is loaded into a process, while system calls allow for code in other processes or the OS kernel to be accessed.
You may want to refer to the background from [[Operating Systems 2026F: Tutorial 1|Tutorial 1]] as well.


===Online Documentation (man pages)===
===Online Documentation (man pages)===
Line 68: Line 70:
Note that <tt>whatis</tt> gives you a man summary, and <tt>apropos</tt> is a quick way to search man pages.
Note that <tt>whatis</tt> gives you a man summary, and <tt>apropos</tt> is a quick way to search man pages.


===The Shell===


The shell or command line provides a text interface for running programs.  While not as visually pleasing as a graphical interface, the shell provides a more clear representation of the functionality provided by the operating system.
===Downloading Code & Compiling Programs===
 
To download C programs to your VM, use <tt>wget</tt> or <tt>curl</tt> commands:
 
wget https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c
  curl https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c -o hello.c
 
To compile, use gcc:
 
gcc -O2 hello.c -o hello
 
This compiles it with level 2 optimization and without debugging symbols.


To run a program contained in the current directory in the shell, you need to prefix the name of the command with a <tt>./</tt>.  This "./" tells the shell that the location of the command you wish to run is the current directory.  By default, the shell will not search for executable commands in the current working directory.  To run most system commands, the name of the command can be typed without a path specification.
To run, you have to specify where it is:


Note that there are many kinds of shells, and people can be very opinionated about which shell is best.  We will be normally using <tt>bash</tt>, but there are many others including ones that have been around forever (sh, csh, tcsh) and somewhat newer, more feature-filled shells (ksh and zsh). There are also shells that were first built for non-UNIX-like systems but now run on Linux (Powershell). Wikipedia has a nice article [https://en.wikipedia.org/wiki/Comparison_of_command_shells comparing the features of different command shells].
  ./hello


===Shell Basics===
Remember you can change directories using the <tt>cd</tt> command.


Note that <tt>bash</tt> is the default shell on most Linux systems.  Other UNIX-like systems can default to other shells like <tt>csh</tt> or <tt>tcsh</tt>; there are many alternatives such as <tt>zsh</tt> that you may prefer.  When you change shells the syntax of the following operations can change; however, conceptually all UNIX-like shells provide the same basic functionality:
By default gcc-produced binaries are dynamically linked (so at runtime they will require dynamic libraries to be present on the system).  To compile a binary that is statically linked (so it has no external runtime library dependencies), instead do this:
* run external programs with command-line arguments
* view and set environment variables
* redirect program input and output using I/O redirection and pipes.
* allow for the creation of scripts that combine external programs with built-in programming functionality.


===Processes===
gcc -O2 -static hello.c -o hello


Each application running on a system is assigned a unique process identifier.  The <tt>ps</tt> command shows the process identifiers for running processes.  Each process running on the system is kept separated from other processes by the operating system.  This information will be useful for subsequent questions.
===Assembly Language===


When you enter a command at a shell prompt, most of the time you are creating a new process which runs the program you specified.
When we write code in C, it has to be compiled to machine code before it can be run.  This compilation step doesn't happen all at once. Compilation has distinct phases:


===Permissions===
* Compile C code into assembly code (.s files).
* Assemble assembly code into machine code placed in object code files (.o files).
* Link object code files together to create a runnable binary.


Your permission to access files in Unix is determined by who you are logged in.  A logged in user has a user ID and belongs to one or more groups.
If you run <tt>gcc -v</tt>, you'll see these steps all happen in a very verbose fashion.


A file is always owned by someone and is always associated with a group.  All files on the Unix file system (including directories and other special files) have three different sets of permissions:
===Static & Dynamic Libraries===
* owner permissions
* group permissions
* other permissions
Each of these have read, write, and/or execute permissions along with some other special permissions we'll discuss later.


The <tt>ls</tt> command with the <tt>-l</tt> option can be used to show both the permissions of a file as well as the owner and group associated with the file.  Permissions are listed first, followed by the owner and the group.
Most applications are not self contained; they rely on lots of external code.  In compiled languages such as C, external code can be brought into the process through '''linking'''.  There are two basic types of linking, static and dynamic linking:
* With '''static linking''', code is brought in at compile time (specifically, in the link stage) and added to the executable.  The code is now the same as other application code.  (Static libraries are just collections of .o files.)
* With '''dynamic linking''', a reference to the library code is added to the binary.  The actual library code has to later be loaded when the program is executed.  This loading will happen before <tt>main()</tt> is called.
Static and dynamic libraries are stored in /lib and /usr/lib, traditionally.


===Environment & Shell Variables===
The dynamic libraries associated with a program binary can be found using the <tt>ldd</tt> command.  You can use <tt>ltrace</tt> to see calls to functions that are dynamically linked.  If a program is statically linked <tt>ldd</tt> will have nothing to report and will generally produce an error.


Environment variables on both Linux and Windows are variable-value pairs that are shared between processes that define important context-related information (such as the name of the current user, the current language, the timezone) for applications.  The key advantage of environment variables is that they are available right when a program starts - they are given to it by the operating system.
Note that code in static and dynamic libraries runs inside of the process loading the code; thus, library code has the same privileges as other application code.  It can do everything regular application code can do (it can access all of your code and data), but it can do no more than your code can do (it has the same restrictions on accessing system resources).


In Linux, these environment variables can be printed on the command line in most shells by referring to the variable name prefixed with a $ sign (eg: to output the value in the HELLO environment variable, one could write <tt>echo $HELLO</tt>).
===System Calls===


Most shells also have internal variables which are private to the shell process.  Typically you can access shell and environment variables using the same mechanisms.  By convention, shell variables are lower case or mixed case, while environment variables are all upper case.  In bash, by default all variables are first shell variables.  To make them environment variables, they must be "export"-ed.  Thus
A process on its own has limited access to the system.  It cannot directly access any external devices or data sources (e.g., files, keyboard, the screen, networks) on its own.  To access these external resources, to allocate memory, or otherwise change its runtime environment, it must make '''system calls'''.  Note that system calls run code outside of a process and thus cannot be called like regular function calls.  The standard C library provides function wrappers for most commonly-used system calls so they can be accessed like regular C functions.  Under the hood, however, these functions make use of special compiler directives in order to generate the machine code necessary to invoke system calls.


  X="Important Data"
You can see the system calls produced by a process using the <tt>strace</tt> command.


just defines X for the current bash process.  However, if you then type
In general, the code you call through a system call has more privileges than regular application code.  This is because a system call is a request to the operating system kernel to do something on behalf of the process, and the kernel has full privileges to the system.  (Indeed, it is the part of the system that implements the process abstraction.) We're going to talk a lot about system calls this semester, this is just your introduction to the concept.


  export X
==Tasks==


X will be turned into an environment variable, and so every subsequent program will also get X.  You can combine both in one line:
===A: Exploring the Openstack VM===


  export X="Important Data"
<ol>
<li>Join the class Team if you haven't already. Link is in brightspace.</li>
<li>How does the environment of the Openstack VM compare to the LinuxOnTab one from Tutorial 1? Specifically, how do the following compare?</li>
<ol style="list-style-type:lower-alpha">
<li>The version of your Linux distribution and the version of your Linux kernel.</li>
<li>The name (binary path) of the current shell, and the shell version.</li>
<li>RAM, disk space, and CPU.</li>
</ol>
<li>Using the man command, find out what the following commands do: <tt>which</tt>, <tt>pwd</tt>, <tt>who</tt>, <tt>whoami</tt>, <tt>env</tt> and
<tt>whereis</tt>.  Try using each of them.</li>
<li>Does the Openstack VM run the same or different programs for standard Linux commands? How do you know?</li>
<li>Making your own commands: the PATH environment variable lists the directories the shell uses to search for external commands. Where can you find documentation on it? How can you add the current directory (whichever directory you are currently in) to PATH? Then, how to make that change permanent? Try to identify multiple ways.</li>
<li>Compile and run [http://people.scs.carleton.ca/~soma/os-2019w/code/csimpleshell.c <tt>csimpleshell.c</tt>]. How does its functionality compare to that of bash? List at least 3 differences.</li>
</ol>


This is the idiom for setting environment variables normally.
===B: Function calls, library calls, and system calls===


To delete an environment variable, you can unset X.
For [http://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c hello.c] and [http://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/syscall-hello.c syscall-hello.c] do the following (<b>substituting the appropriate source file for prog.c</b>). For example, for hello.c, you would replace all instances of "prog" in a command with "hello".


One thing to remember with the above is that spaces are used to separate arguments in bash and most other UNIX shells.  Thus it is an error to type:
To download programs to your VM, use the wget command, e.g.
<pre>
wget https://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c
</pre>
# Compile the program prog.c using <tt>gcc -O2 prog.c -o prog-dyn</tt> and run prog-dyn.  What does it do?
# Statically compile and optimize prog.c by running <tt>gcc -O2 -static prog.c -o prog-static</tt>.  How does the size compare with <tt>prog</tt>?
# Run <tt>ldd</tt> on the static and dynamic versions of the program.  How does the output compare?  Why?
# See what system calls prog-static produces by running <tt>strace -o syscalls-static.log ./prog-static</tt>.  Do the same for <tt>prog-dyn</tt>.  Which version generates more system calls?  '''Note: system calls are saved in the log file syscalls-static.log.  Feel free to save them in a different file.'''
# See what library calls prog-static produces by running <tt>ltrace -x '*' -o library-static.log ./prog-static</tt>.  Do the same for <tt>prog-dyn</tt>.  Which version generates more library calls?  (Note: you will have to run <tt>sudo apt install ltrace</tt> to enable the command.)
# Compile the program dynamically but with "lazy" linking: <tt>gcc -O2 -z lazy prog.c -o prog-dynlazy</tt>.  Run <tt>ltrace -o library-lazy.log ./prog-dynlazy</tt>.  How does the output of this compare to that of the previous ltrace?
# (optional) Look up the documentation for each of the system calls made by the static versions of the programs.  You may need to append a 2 or 3 to the manpage invocation, e.g. "man 2 write" gets you the write system call documentation.


  export X = "Important Data"
===C: Comparing C and assembly===


as you now are giving export three arguments, not one.
Do the following with hello.c and syscall-hello.c, as before.


One of the key reasons people choose alternatives to bash is because of quirks like this!
A few tips on x86-64 assembly language:
* The last letter of many instructions refers to the size of the operand.  For example, callq means call a function using a "quad" value (64 bits).
* A dollar sign preceding a value means that it is a literal value, a percent sign means it is a register.
* If a register is in parentheses, then it is being used as a "pointer" (it contains an address, so the CPU goes to that address and interacts with the memory there).  If there is a number before the parentheses, it is an offset to the register's value.


===Controlling Processes===
Resources on x86-64 assembly language (only needed if you want to learn more):
* [https://www.cs.cmu.edu/~fp/courses/15213-s07/misc/asm64-handout.pdf CMU introduction to x86-64]
* [https://en.wikibooks.org/wiki/X86_Assembly/GAS_Syntax AT&T/GNU Assembler syntax]
* [https://en.wikipedia.org/wiki/X86_calling_conventions#System_V_AMD64_ABI the Wikipedia article on calling conventions]


On Linux, you can control processes by sending them signals.
# Using the <tt>nm</tt> command, see what symbols are defined in prog-static and prog-dyn.  Which defines more symbols?
# Run the command <tt>gcc -c -O2 prog.c</tt> to produce an object file.  What file was produced?  What symbols does it define?
# Look at the assembly code of the program by running <tt>gcc -S -O2 prog.c</tt>.  What file was produced?  Identify the following in the assembly code (if present):
#* A function call (call)
#* A return from a function (ret)
#* Registers being saved onto the stack (push)
#* Registers being retrieved from the stack (pop)
#* Subtraction (sub)
#* A system call (syscall)
# Disassemble the object file using <tt>objdump -d</tt>.  How does this disassembly compare with the output from gcc -S?
# Examine the headers of object file, dynamically linked executable, and the statically linked executable using <tt>objdump -h</tt>
# Examine the contents of object file, dynamically linked executable, and the statically linked executable using <tt>objdump -s</tt>
# Re-run all of the previous gcc commands adding the "-v" flag. What is all of that output?


You send signals when you type certain key sequences in most shells: Control-C sends INT (interrupt), Control-Z sends STOP.
===D: Examining the runtime memory map===


You can send a signal to a process using the kill command:
Compile and run [https://homeostasis.scs.carleton.ca/~soma/os-2019f/code/3000memview.c 3000memview.c], then consider the following questions.
# Why are the addresses inconsistent between runs?  What happens if you run the program with the command "setarch -R ./3000memview"?  (Note this changes how 3000memview is run.)
# Roughly where does the stack seem to be?  The heap?  Code?  Global variables?
# Observe how the heap grows (i.e. the value of sbrk changes) in response to malloc calls.  Would you expect the heap to ever run into the stack?  Why or why not?
# Change each malloc call to allocate more than 128K.  What happens to the values of sbrk?  Why?  (Hint: use strace)
# Add more code and data to the program, and add more printf's to see where things are.  Are things where you expect them to be?


  kill -<signal> <process ID>
==Code==


So to stop process 4542, type
===[https://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c hello.c]===
<syntaxhighlight lang="c" line>
#include <stdio.h>


  kill -STOP 4542
int main(int argc, char *argv[]) {


By default, kill sends the TERM signal.
        printf("Hello world!\n");


===Downloading Code & Compiling Programs===
        return 0;
}
</syntaxhighlight>


To download C programs to your VM, use <tt>wget</tt> or <tt>curl</tt> commands:
===[http://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/syscall-hello.c syscall-hello.c]===
<syntaxhighlight lang="c" line>
#include <unistd.h>
#include <sys/syscall.h>


wget https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c
char *buf = "Hello world!\n";
curl https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c -o hello.c


To compile, use gcc:
int main(int argc, char *argv) {
        size_t result;


gcc -O2 hello.c -o hello
        /* "man 2 write" to see arguments to write syscall */
        result = syscall(SYS_write, 1, buf, 13);


This compiles it with level 2 optimization and without debugging symbols.
        return (int) result;
}
</syntaxhighlight>


To run, you have to specify where it is:
===[https://homeostasis.scs.carleton.ca/~soma/os-2019f/code/3000memview.c 3000memview.c]===
<syntaxhighlight lang="c" line>
#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>


./hello
char *gmsg = "Global Message";


Remember you can change directories using the <tt>cd</tt> command.
const int buffer_size = 100;


By default gcc-produced binaries are dynamically linked (so at runtime they will require dynamic libraries to be present on the system).  To compile a binary that is statically linked (so it has no external runtime library dependencies), instead do this:
int main(int argc, char *argv[], char *envp[])
{
        char *lmsg = "Local Message";
        char *buf[buffer_size];
        int i;
       
        printf("Memory report\n");
        printf("argv:      %lx\n", (unsigned long) argv);
        printf("argv[0]:  %lx\n", (unsigned long) argv[0]);
        printf("envp:      %lx\n", (unsigned long) envp);
        printf("envp[0]:   %lx\n", (unsigned long) envp[0]);


gcc -O2 -static hello.c -o hello
        printf("lmsg:      %lx\n", (unsigned long) lmsg);
        printf("&lmsg:    %lx\n", (unsigned long) &lmsg);
        printf("gmsg:      %lx\n", (unsigned long) gmsg);
        printf("&gmsg:    %lx\n", (unsigned long) &gmsg);


==Tasks/Questions==
        printf("main:      %lx\n", (unsigned long) &main);


<ol>
        printf("sbrk(0):   %lx\n", (unsigned long) sbrk(0));
<li>When you have logged in to a shell, how (i.e., using what commands?) do you first find out information about the environment?</li>
        printf("&buf:     %lx\n", (unsigned long) &buf);
<ol style="list-style-type:lower-alpha">
<li>The version of your Linux distribution and the version of your Linux kernel.</li>
<li>The name (binary path) of the current shell.</li>
<li>RAM, disk space, and CPU.</li>
</ol>
<li>Using the man command, find out what the following commands do: <tt>which</tt>, <tt>pwd</tt>, <tt>who</tt>, <tt>whoami</tt>, <tt>env</tt> and
<tt>whereis</tt>.  Try using each of them.</li>
<li>Linux commands can be classified as internal (built into the shell) and external (separate program binaries). How can you tell if a specific command (e.g., cd) is internal or external? Figure out where at least three external commands reside on the system.</li>
<li>Making your own commands: the PATH environment variable lists the directories the shell uses to search for external commands. Where can you find documentation on it? How can you add the current directory (whichever directory you are currently in) to PATH? Then, how to make that change permanent? Try to identify multiple ways.</li>
<li>Look at the permissions of the program binaries of the external commands you have just found above. Who owns them? What group are they in?</li>
<li>For those same program binaries, figure out what the permission bits mean by reading the man page of chmod (this is the command you could use to change those permission bits).</li>
<li>What are the owner, group, and permissions of /etc/passwd and /etc/shadow? What are these files used for?</li>
<li>What does it mean to have execute permission on a directory?</li>
<li>The <tt>ls</tt> command can be used to get a listing of the files in a directory. What options are passed to <tt>ls</tt> to see: the permission bits above; all the files within a directory (including hidden files)? How
to make a file hidden?</tt>
<li>Compile and run [http://people.scs.carleton.ca/~soma/os-2019w/code/csimpleshell.c <tt>csimpleshell.c</tt>]. How does its functionality compare to that of bash? List at least 3 differences.</li>
</ol>


==Code==
        for (i = 0; i<buffer_size; i++) {
                buf[i] = (char *) malloc(4096);
        }
       
        printf("buf[0]:    %lx\n", (unsigned long) buf[0]);
        printf("sbrk(0):  %lx\n", (unsigned long) sbrk(0));
       
        return 0;
}
</syntaxhighlight>


===[http://people.scs.carleton.ca/~soma/os-2019w/code/csimpleshell.c csimpleshell.c]===
===[http://people.scs.carleton.ca/~soma/os-2019w/code/csimpleshell.c csimpleshell.c]===

Latest revision as of 21:14, 24 September 2026

In this tutorial we're going to look at how processes work at a low level: how they make system calls & library calls, how C and assembly compare, and and how memory is laid out.

Getting Started

For this tutorial, you need to get access to a Linux or UNIX machine, and LinuxOnTab isn't enough. We strongly suggest you use an SCS Openstack instance (see below). You'll need access to a system for the entire semester, ideally the same one.

The concepts covered below are mostly part of standard UNIX/Linux tutorials. Feel free to consult one or more of them. However, remember that you are trying to build a conceptual model of how things work. Thus, don't memorize commands; instead, try to understand how things fit together, and ask questions when things don't work as expected!

If you find yourself searching for the answers to specific questions, you're probably doing it wrong.

Feel free to discuss this tutorial on Teams in the Tutorials channel.

Again, for emphasis: don't take snapshots! (see below)

Connecting to Openstack

SCS has lots of documentation on openstack, including a step-by-step guide. Start here!

Create a VM on the new SCS openstack cluster at openstack-stein.scs.carleton.ca and do your work there. While you don't need a persistent VM for this lab, it will be important for future tuturials - and it is nice to have your work stick around when you leave the lab.

To access Openstack you must be on the Carleton network, so make sure to VPN in if you aren't on the Carleton computer or WiFi network.

To get added to the COMP 3000 project, you need to change your SCS password (run newacct) in order to update your account to have the right entitlements.

Setting up and connecting a VM

Create a VM on the SCS openstack cluster as shown in the SCS documentation. Make sure do the following:

  • Choose the COMP3000-F26.2026-* instance snapshot image ONLY (others won't have the right software for later tutorials), use the latest one.
  • Add the ping-ssh-egress security group, and
  • Associating a floating IP address.

The 192.168.X.X IP addresses are private (and cannot be accessed outside of the openstack cluster), the 134.117.X.X floating IP addresses can be accessed from the Carleton network and will allow you to access the wider Internet. You need to ssh to your VM instance. Windows, Ubuntu and MacOS all have SSH clients available from their command lines, just type "ssh student@<IP address>" where the IP address is the floating IP address you assigned to your VM (while connected to the Carleton VPN). Other tools supporting SSH (e.g., PuTTY) also work.

Once you are prompted to log in, the default user is student, default password is student. You'll have to change your password after you first login. (If you want to change your password later, use the passwd command.)

You can also connect directly to your instance with ssh -J (proxy), going through access:

 ssh -J <SCS username>@access.scs.carleton.ca student@<Openstack floating IP address>

(Don't use the web console unless it is an emergency, it will be glitchy. Also, x2go won't work because the VM doesn't have a desktop environment installed, on purpose.)

Backups

The image provides an "scs-backup" command that will backup the student user's directory to the SCS linux machines. So if your SCS username is janedoe, you can type:

scs-backup janedoe

and it will create a copy of everything (note: you can customize it) in the student account in a directory called "COMP3000VM-backup" in your home directory. You can ssh/sftp to access.scs.carleton.ca in order to access this copy of your VM's files.

You should do backups at the end of every session and before you do anything dangerous. While the cluster is generally stable, you should be ready for everything in it to be erased at a moment's notice, because it could happen!

Note that you cannot take snapshots of your VM, so please don't try (it will keep trying and never succeed, and you'll make work for the tech staff who have to cancel what you did).

Background

A process on UNIX-like systems are separated from each other: they run in their own address space - pointers can only refer to code and data in that program, not in other programs. Because programs are separated, they can't access anything external to them without help.

Here we're going to talk about libraries and system calls, two ways programs gain access to additional functionality. Libraries are for external code that is loaded into a process, while system calls allow for code in other processes or the OS kernel to be accessed.

You may want to refer to the background from Tutorial 1 as well.

Online Documentation (man pages)

The man (short for manual) command is one of the primary ways to to access built-in software documentation. Most software packages that provide command-line programs include man pages.

For almost any commands mentioned in the tutorials, you can use man to find the usage. While you can also find documentation online for these same commands, many have multiple variants that have different functionality. The man page is guaranteed to document the version installed on your system.

Man pages are divided into multiple sections, with each section having its own purpose, e.g., 1 for general commands, 2 for system calls, and 3 for library functions. You can specify the section as the first argument to man if there is more than one man page with the same name. For instance, tee is both a command (man 1 tee) and a system call (man 2 tee). The lowest number man page will be returned if the section is not specified.

Note that the topics in man pages go beyond just software & command manuals; they also include conventions and abstract concepts (e.g., man syscalls and man man-pages). Thus if you have questions, consider browsing the man pages rather than just going to a search engine.

Note that whatis gives you a man summary, and apropos is a quick way to search man pages.


Downloading Code & Compiling Programs

To download C programs to your VM, use wget or curl commands:

wget https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c
curl https://people.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c -o hello.c

To compile, use gcc:

gcc -O2 hello.c -o hello

This compiles it with level 2 optimization and without debugging symbols.

To run, you have to specify where it is:

./hello

Remember you can change directories using the cd command.

By default gcc-produced binaries are dynamically linked (so at runtime they will require dynamic libraries to be present on the system). To compile a binary that is statically linked (so it has no external runtime library dependencies), instead do this:

gcc -O2 -static hello.c -o hello

Assembly Language

When we write code in C, it has to be compiled to machine code before it can be run. This compilation step doesn't happen all at once. Compilation has distinct phases:

  • Compile C code into assembly code (.s files).
  • Assemble assembly code into machine code placed in object code files (.o files).
  • Link object code files together to create a runnable binary.

If you run gcc -v, you'll see these steps all happen in a very verbose fashion.

Static & Dynamic Libraries

Most applications are not self contained; they rely on lots of external code. In compiled languages such as C, external code can be brought into the process through linking. There are two basic types of linking, static and dynamic linking:

  • With static linking, code is brought in at compile time (specifically, in the link stage) and added to the executable. The code is now the same as other application code. (Static libraries are just collections of .o files.)
  • With dynamic linking, a reference to the library code is added to the binary. The actual library code has to later be loaded when the program is executed. This loading will happen before main() is called.

Static and dynamic libraries are stored in /lib and /usr/lib, traditionally.

The dynamic libraries associated with a program binary can be found using the ldd command. You can use ltrace to see calls to functions that are dynamically linked. If a program is statically linked ldd will have nothing to report and will generally produce an error.

Note that code in static and dynamic libraries runs inside of the process loading the code; thus, library code has the same privileges as other application code. It can do everything regular application code can do (it can access all of your code and data), but it can do no more than your code can do (it has the same restrictions on accessing system resources).

System Calls

A process on its own has limited access to the system. It cannot directly access any external devices or data sources (e.g., files, keyboard, the screen, networks) on its own. To access these external resources, to allocate memory, or otherwise change its runtime environment, it must make system calls. Note that system calls run code outside of a process and thus cannot be called like regular function calls. The standard C library provides function wrappers for most commonly-used system calls so they can be accessed like regular C functions. Under the hood, however, these functions make use of special compiler directives in order to generate the machine code necessary to invoke system calls.

You can see the system calls produced by a process using the strace command.

In general, the code you call through a system call has more privileges than regular application code. This is because a system call is a request to the operating system kernel to do something on behalf of the process, and the kernel has full privileges to the system. (Indeed, it is the part of the system that implements the process abstraction.) We're going to talk a lot about system calls this semester, this is just your introduction to the concept.

Tasks

A: Exploring the Openstack VM

  1. Join the class Team if you haven't already. Link is in brightspace.
  2. How does the environment of the Openstack VM compare to the LinuxOnTab one from Tutorial 1? Specifically, how do the following compare?
    1. The version of your Linux distribution and the version of your Linux kernel.
    2. The name (binary path) of the current shell, and the shell version.
    3. RAM, disk space, and CPU.
  3. Using the man command, find out what the following commands do: which, pwd, who, whoami, env and whereis. Try using each of them.
  4. Does the Openstack VM run the same or different programs for standard Linux commands? How do you know?
  5. Making your own commands: the PATH environment variable lists the directories the shell uses to search for external commands. Where can you find documentation on it? How can you add the current directory (whichever directory you are currently in) to PATH? Then, how to make that change permanent? Try to identify multiple ways.
  6. Compile and run csimpleshell.c. How does its functionality compare to that of bash? List at least 3 differences.

B: Function calls, library calls, and system calls

For hello.c and syscall-hello.c do the following (substituting the appropriate source file for prog.c). For example, for hello.c, you would replace all instances of "prog" in a command with "hello".

To download programs to your VM, use the wget command, e.g.

 wget https://homeostasis.scs.carleton.ca/~soma/os-2017f/code/tut1/hello.c
  1. Compile the program prog.c using gcc -O2 prog.c -o prog-dyn and run prog-dyn. What does it do?
  2. Statically compile and optimize prog.c by running gcc -O2 -static prog.c -o prog-static. How does the size compare with prog?
  3. Run ldd on the static and dynamic versions of the program. How does the output compare? Why?
  4. See what system calls prog-static produces by running strace -o syscalls-static.log ./prog-static. Do the same for prog-dyn. Which version generates more system calls? Note: system calls are saved in the log file syscalls-static.log. Feel free to save them in a different file.
  5. See what library calls prog-static produces by running ltrace -x '*' -o library-static.log ./prog-static. Do the same for prog-dyn. Which version generates more library calls? (Note: you will have to run sudo apt install ltrace to enable the command.)
  6. Compile the program dynamically but with "lazy" linking: gcc -O2 -z lazy prog.c -o prog-dynlazy. Run ltrace -o library-lazy.log ./prog-dynlazy. How does the output of this compare to that of the previous ltrace?
  7. (optional) Look up the documentation for each of the system calls made by the static versions of the programs. You may need to append a 2 or 3 to the manpage invocation, e.g. "man 2 write" gets you the write system call documentation.

C: Comparing C and assembly

Do the following with hello.c and syscall-hello.c, as before.

A few tips on x86-64 assembly language:

  • The last letter of many instructions refers to the size of the operand. For example, callq means call a function using a "quad" value (64 bits).
  • A dollar sign preceding a value means that it is a literal value, a percent sign means it is a register.
  • If a register is in parentheses, then it is being used as a "pointer" (it contains an address, so the CPU goes to that address and interacts with the memory there). If there is a number before the parentheses, it is an offset to the register's value.

Resources on x86-64 assembly language (only needed if you want to learn more):

  1. Using the nm command, see what symbols are defined in prog-static and prog-dyn. Which defines more symbols?
  2. Run the command gcc -c -O2 prog.c to produce an object file. What file was produced? What symbols does it define?
  3. Look at the assembly code of the program by running gcc -S -O2 prog.c. What file was produced? Identify the following in the assembly code (if present):
    • A function call (call)
    • A return from a function (ret)
    • Registers being saved onto the stack (push)
    • Registers being retrieved from the stack (pop)
    • Subtraction (sub)
    • A system call (syscall)
  4. Disassemble the object file using objdump -d. How does this disassembly compare with the output from gcc -S?
  5. Examine the headers of object file, dynamically linked executable, and the statically linked executable using objdump -h
  6. Examine the contents of object file, dynamically linked executable, and the statically linked executable using objdump -s
  7. Re-run all of the previous gcc commands adding the "-v" flag. What is all of that output?

D: Examining the runtime memory map

Compile and run 3000memview.c, then consider the following questions.

  1. Why are the addresses inconsistent between runs? What happens if you run the program with the command "setarch -R ./3000memview"? (Note this changes how 3000memview is run.)
  2. Roughly where does the stack seem to be? The heap? Code? Global variables?
  3. Observe how the heap grows (i.e. the value of sbrk changes) in response to malloc calls. Would you expect the heap to ever run into the stack? Why or why not?
  4. Change each malloc call to allocate more than 128K. What happens to the values of sbrk? Why? (Hint: use strace)
  5. Add more code and data to the program, and add more printf's to see where things are. Are things where you expect them to be?

Code

hello.c

#include <stdio.h>

int main(int argc, char *argv[]) {

        printf("Hello world!\n");

        return 0;
}

syscall-hello.c

#include <unistd.h>
#include <sys/syscall.h>

char *buf = "Hello world!\n";

int main(int argc, char *argv) {
        size_t result;

        /* "man 2 write" to see arguments to write syscall */
        result = syscall(SYS_write, 1, buf, 13);

        return (int) result;
}

3000memview.c

#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>

char *gmsg = "Global Message";

const int buffer_size = 100;

int main(int argc, char *argv[], char *envp[])
{
        char *lmsg = "Local Message";
        char *buf[buffer_size];
        int i;
        
        printf("Memory report\n");
        printf("argv:      %lx\n", (unsigned long) argv);
        printf("argv[0]:   %lx\n", (unsigned long) argv[0]);
        printf("envp:      %lx\n", (unsigned long) envp);
        printf("envp[0]:   %lx\n", (unsigned long) envp[0]);

        printf("lmsg:      %lx\n", (unsigned long) lmsg);
        printf("&lmsg:     %lx\n", (unsigned long) &lmsg);
        printf("gmsg:      %lx\n", (unsigned long) gmsg);
        printf("&gmsg:     %lx\n", (unsigned long) &gmsg);

        printf("main:      %lx\n", (unsigned long) &main);

        printf("sbrk(0):   %lx\n", (unsigned long) sbrk(0));
        printf("&buf:      %lx\n", (unsigned long) &buf);

        for (i = 0; i<buffer_size; i++) {
                buf[i] = (char *) malloc(4096);
        }
        
        printf("buf[0]:    %lx\n", (unsigned long) buf[0]);
        printf("sbrk(0):   %lx\n", (unsigned long) sbrk(0));
        
        return 0;
}

csimpleshell.c

/* csimpleshell.c, Enrico Franchi © 2005
      https://web.archive.org/web/20170223203852/
      http://rik0.altervista.org/snippets/csimpleshell.html
      "BSD" license

   January 12, 2019: minor changes to eliminate most compilation warnings
   (Anil Somayaji, soma@scs.carleton.ca)
*/

#include <stdio.h>
#include <stdlib.h>
#include <unistd.h>
#include <string.h>
#include <errno.h>
#include <sys/wait.h>
#include <sys/types.h>
#define BUFFER_SIZE 1<<16
#define ARR_SIZE 1<<16

void parse_args(char *buffer, char** args, 
                size_t args_size, size_t *nargs)
{
    char *buf_args[args_size]; /* You need C99 */
    char **cp;
    char *wbuf;
    size_t i, j;
    
    wbuf=buffer;
    buf_args[0]=buffer; 
    args[0] =buffer;
    
    for(cp=buf_args; (*cp=strsep(&wbuf, " \n\t")) != NULL ;){
        if ((*cp != NULL) && (++cp >= &buf_args[args_size]))
            break;
    }
    
    for (j=i=0; buf_args[i]!=NULL; i++){
        if(strlen(buf_args[i])>0)
            args[j++]=buf_args[i];
    }
    
    *nargs=j;
    args[j]=NULL;
}


int main(int argc, char *argv[], char *envp[]){
    char buffer[BUFFER_SIZE];
    char *args[ARR_SIZE];

    int ret_status;
    size_t nargs;
    pid_t pid;
    
    while(1){
        printf("$ ");
        fgets(buffer, BUFFER_SIZE, stdin);
        parse_args(buffer, args, ARR_SIZE, &nargs); 

        if (nargs==0) continue;
        if (!strcmp(args[0], "exit" )) exit(0);       
        pid = fork();
        if (pid){
            printf("Waiting for child (%d)\n", pid);
            pid = wait(&ret_status);
            printf("Child (%d) finished\n", pid);
        } else {
            if( execvp(args[0], args)) {
                puts(strerror(errno));
                exit(127);
            }
        }
    }    
    return 0;
}