Wednesday, September 14, 2011

Mathematical Induction

I actually posted on this very topic a few years ago.  However, I have refined how I think about doing proofs by mathematical induction.  And so I am writing one more time.

First, you might find it interesting to look at the Wikipedia article on mathematical induction.

The Principle of Mathematical Induction (PMI) is an axiom that describes the set of natural numbers, which is N = {1, 2, 3, 4, ...}.  From one point of view, the PMI says that if S is a set that (1) contains the number 1 and (2) guarantees that whenever any number n is in the set, then n+1 is also in the set, then the set S contains all of N.  From the perspective of proofs, however, we are really interested in showing that an infinite chain of logical statements is true.

Before we proceed with the idea of the PMI, let me be clear about a logical statement.  A logical statement is a well-constructed sentence that is definitively true or false. Two well-known but easily misunderstood examples of logical statements are equations and inequalities.  An equation "A=B" is a logical statement that declares two quantities (A and B) are equal (have the same value).  An inequality "Afalse, but the statement itself is logically complete.  (Here logical does not mean `makes sense' but it means `can be decided between True or False'.)

Related to this issue is a common misunderstanding by students that "=" is like an operation that means "has the value" or "find the answer" as it is often used on a calculator.  This is especially important in a proof, where each statement (usually an equation) must be clearly true rather than a statement of what you hope or what you are trying to calculate.

Okay, back to the Principle of Mathematical Induction.  This principle is about dealing with an infinite chain of logical sentences.  For example, consider the following chain of statements.  S(1) is the label for the first sentence, S(2) is the label for the second sentence, and so on:

S(1):     1 = 1(2)/2
S(2):     1+2 = 2(3)/2
S(3):     1+2+3 = 3(4)/2
S(4):     1+2+3+4 = 4(5)/2
S(5):     1+2+3+4+5 = 5(6)/2
S(6):     1+2+3+4+5+6 = 6(7)/2
...
(the pattern continues so that if n is any number n=1,2,3,..., we have):
S(n):     1+2+3+...+n = n(n+1)/2

On the left of each sentence is a summation.  On the right of each sentence is a simple formula involving n.  Using summation notation, we would have written this sentence:
S(n):    Σnk=1 k = n(n+1)/2

The pattern described by S(n) looks like a single sentence, but it is important to remember that it is describing the entire infinitely long chain of logical sentences.  If you add up the values on the left side of any single sentence and compare it to the value of the formula on the right, you will see that the answers are the same.  But this only verifies the formulas that you actually check.

The Principle of Mathematical Induction is a tool to prove that the entire chain of sentences are all true.  The PMI is much like an infinite chain of dominoes (see the earlier linked wikipedia article).  To knock over a chain, if you knock them down one at a time, you'll never finish.  But if you show that knocking any single domino down will guarantee the next domino also falls and you show that you can knock down the first domino, then this is enough to guarantee that the entire chain will fall down.

So here is the PMI:
Suppose S(n), n=1, 2, 3, ..., is a chain of logical statements and that (1) S(1) is true and (2) S(n) implies S(n+1) for any n=1, 2, 3, ..., then S(n) is true for every n=1, 2, 3, ....
Notice that there are two conditions to use PMI:
  1. We must verify that S(1) (the first sentence) is true.
  2. We must show that the sentence S(n+1) is true whenever the previous sentence S(n) is true.
A proof by induction is the process whereby we verify these two conditions and then apply the PMI to conclude that the entire chain is true.

To learn more about writing the proofs and an example with commentary, please continue by reading the next blog post.

Tuesday, March 8, 2011

How We Learn Mathematics

I was reading the following paper: D. Breidenbach, E. Dubinsky, J. Hawks, and D. Nichols, "Development of the Process Conception of Function," Educational Studies in Mathematics, 23: 247-285, 1992.

Quote Dubinsky (1989): "A person's mathematical knowledge is her or his tendency to respond to certain kinds of perceived problem situations by constructing, reconstructing and organizing mental processes and objects to use in dealing with the situations."

"Applying this point of view to mathematics (or any other subject) consists of determining the nature of the specific processes and objects that are constructed and how they are organized when one studies mathematics"

Ways of thinking about functions:
  • prefunction - does not understand any real ways of using function concepts
  • action - repeatable mental or physical manipulation (e.g., plug in numbers and calculate); static; one step at a time
  • process - think of function as a single dynamic transformation
I then found another article: A. Sfard and L. Linchevski, "The Gains and the Pitfalls of Reification: The Case of Algebra," Educational Studies in Mathematics, 26 (2/3), 191-228, 1994 [Learning Mathematics: Constructivist and Interactionist Theories of Mathematical Development]

This article proceeds with the view that in mathematics, there is a duality in mathematical constructs being a process or an object. That is, conceive of things operationally (process) or structurally (object). Historical examples include the expansion of number systems: positive to negative (operational: subtraction as adding a negative to structural: negative numbers as objects), and real to complex (i=sqrt(-1) as an operational convenience to an actual object)

An included reference suggests finding another article: Kieran, C.: 1992, 'The learning and teaching of school algebra', in D. A. Grouws (ed.), The Handbook of Research on Mathematics Teaching and Learning, Macmillan, New York, pp. 390-419. I'll have to see if I can find this one, as it is cited for the sentence, "[reification] was also used to introduce some order into the quickly growing bulk of findings about algebraic thinking."

Interesting phrase: "the ability to grasp the structural aspect is not easy to achieve" and "those crucial junctions in the development of mathematics where a transition from one level to another takes place are the most problematic."

Another interesting way to think about how mathematics is organized: (1) Logical, or the way it fits together; (2) Historical, or the way in which it was developed; and (3) Cognitive, or the processes in which people learn.

Modes of Algebra
1.1) Algebra as Generalized Arithmetic: The Operational Phase
-- solve for the unknown, but not using symbols (grade school algebraic thinking)
-- rhetoric algebra
-- principally reversing processes
1.2) Algebra as Generalized Arithmetic: The Structural Phase
1.2.1) algebra of a fixed value (unknown)
-- Notational convenience, but treat variable as a fixed value
:::: becomes a mental challenge to think of formula as both a process and result
:::: example given: 2+3 represents process, 5 represents result. But x+3 represents both, no separate "result"
:::: compare to the challenges of new number types required to think about division, subtraction, and extracting square roots
** Nice comment: "Once we manage to overcome this difficulty, it is quickly forgotten. ... Our eyes are easily blinded by habit and by our own ontological beliefs. Nevertheless, much evidence for the difficulty of reification may also be found in today's classroom, provided those who listen to the students are open-minded enough to grasp the ontological gap between themselves and the less experienced learners."
1.2.2) Functional algebra (of a variable)
:::: View formula as object
:::: Parameters represented as symbols not numbers.
2) Abstract Algebra

Give examples of interview questions. Students at early stages of thinking think about formulas as recipes for computations (process) but do not perceive them as valid objects. "The equality sign is interpreted as a 'do something signal' (Behr et al 1976; Kieran 1981)"

Here's something I see all the time in calculus classes: "It [the = symbol] serves here as a 'run' command. When treated in this way, the equality symbol looses [sic] the basic characteristics of an equivalence predicate: it stops being symmetrical or transitive. Indeed, young children seem to have no qualms about solving word problems with the help of a chain of non-transitive equalities. For instance, when asked 'How many marbles do you have after you win 4 marbles 3 times and 2 marbles 5 times?', the child would often write: 3*4=12+5*2=12+10=22."

Equations of the form 2x-3 = 11 can be interpreted as a formula whose result is 11 (which can be solved by inverse operations); equations of the form 2x-3=5x-9 appear to be two different formulas, and inverse operations do not make sense.

Wednesday, April 28, 2010

Graph from a graph of f ' (x)

First, pay attention: the graph provided on the assignment is the graph of the derivative f '(x) and not the graph of f. So you can't look at the picture and say that because the graph you are looking at is increasing that f '(x) is positive; if the graph is increasing, then that means f '(x) is increasing, and not f(x). (This is useful information, but you just need to think about what it does say.)

Second, the number line sign analysis summaries will help identify the shape of the graph. Imagine taking the unit circle and breaking it up according to quadrants. The signs of f '(x) and f ''(x) determine which of these four basic shapes the graph is most like.
  • f '(x) = + and f ''(x) = + means f(x) looks like Quadrant IV (incr, conc. up)
  • f '(x) = - and f ''(x) = + means f(x) looks like Quadrant III (decr, conc. up)
  • f '(x) = + and f ''(x) = - means f(x) looks like Quadrant II (incr, conc. down)
  • f '(x) = - and f ''(x) = - means f(x) looks like Quadrant I (decr, conc. down)
The graph is just formed by taking these shapes and putting them end-to-end. You wouldn't actually use the entire portion of the unit circle because we probably don't want vertical tangents like the unit circle has. The circle just helps us remember the basic shape. The points where we join the shapes together will probably be inflection points (concavity changes) or extreme values.

However, sign analysis does not tell us the heights of any points. The problem gives only one point: f(0) = 1. The rest of the points of interest (especially the local extreme values) can be found by thinking about the information relating to the areas of the graph of f '(x). (Again, think about the Fundamental Theorem of Calculus).

Sums of Geometric Sequences

The second problem on the project introduces a new closed form for a sum:

k=1n [A ρk] = A ρ (ρn-1)/(ρ-1).

Unfortunately, too many of you are still intimidated simply by the symbols that are used.

The formula for the geometric sequence, A ρk, is like an exponential, except the power is an integer variable rather than a continuous variable like x. For example, if A=2 and ρ=1/3, we have terms that are increasing powers of (1/3) times 2:
2/3 (k=1), 2/9 (k=2), 2/27 (k=3), 2/81 (k=4), etc.

The summation is simply the sum of these values:
k=1n [2 (1/3)k] = 2/3 + 2/9 + 2/27 + 2/81 + ... + 2/3n.
The closed form gives a formula answering the value of this sum:
2(1/3)[(1/3)n-1]/[(1/3) - 1] = (2/3)*[(1/3)n-1]/(-2/3) = 1 - (1/3)n.

So when you write down the Riemann sum for the integral in question, you need to look at using the properties of exponentials so that the Riemann sum looks just like a sum of a geometric sequence. You should identify a factor that does not involve k, and this is A. You should identify the other factor as some number raised to the k power. Then you can use the closed form.

Populations, Birth Rates, and Death Rates

This entry is a general assist for my class working on a project. Suppose you knew the rate at which births are occurring (call it a function of time, b(t)) and you knew the rate at which deaths are occurring (a function d(t)). If the only way the population changes is through births and deaths, then if P(t) is the function describing the size of the population in time, then P'(t) = b(t) - d(t). (It is still your job to explain why this makes biological sense.)

Okay, now for the general principle. Anytime you know the rate of change of a quantity, you can always get back to the original quantity through a definite integral (assuming the rate of change is continuous, anyway). This is the heart of the 2nd Fundamental Theorem of Calculus. Not using P and t as variables (so that you have at least something to translate), here is the basic idea.

Suppose you know f '(x). Then A(x) = ∫0x f '(z) dz is an antiderivative of f '. But so is f(x) since that is where f '(x) comes from. That is f(x) = A(x) + C for some constant. In particular, A(0) = 0, so C=f(0). That is, f(x) = f(0) + ∫0x f '(z) dz.

This will always work, even if I don't start the integral at 0: f(x) = f(a) + ∫ax f '(z) dz. Written another way, it looks like the first Fundamental Theorem of Calculus: f(x) - f(a) = ∫ax f '(z) dz.
In other words, the 2nd FTC implies that every function is its starting value plus the integral of its rate of change.

Now, for our population problem, we don't actually know the rate of change completely; we only know the value at specific points. So instead of computing an integral (to get an exact value), we will approximate the integral using a Riemann sum. We are restricted to using the table data, so Δt=2 is forced upon us. For example, ∫02 b(t) dt can only be estimated with a single rectangle while ∫04 b(t) dt would involve two rectangles. The idea of the Riemann sum is that we choose b(tk*) as one of our data points (either on the left or right).

More specifically, on the interval [0,2] (k=1), we can either use t1*=0 so that b(t1*)=100 or use t1*=2 so that b(t1*)=135. In the first case, the rectangle for k=1 contributes b(t1*)Δt = 200, while the second case leads to a contribution of b(t1*)Δt = 270. The average (midpoint) of these two values (200+270)/2 = 235 is the estimate that would come from using the trapezoid sum. We do this for each of the 8 intervals between our data points, for both the births and the deaths.

By considering our estimates for the number of births and deaths in each of the intervals ([0,2], [2,4], [4,6], etc.), we can produce an estimate of the new population at each of the times (2, 4, 6, 8, etc.). By thinking about what estimates lead to the largest predicted population, we get an upper limit (i.e. bound) for our estimate --- no population consistent with this data can ever go above that value. Similarly, we can choose those estimates to create a lower bound. The true population will be somewhere in between.

Friday, February 5, 2010

Exponential Functions

I'm getting feedback that exponential functions are giving you extra trouble. I'd appreciate getting feedback to help know how I can clarify the concepts. Here is a summary of some of the key concepts that I'm wanting you to understand:
  • bx is not just a formula "b to the power x" but is a new function, which I'm asking you to call expb(x).
  • The properties of exponents like bx+y=bx by and (bx)y=bxy become properties of the exponential functions.
    expb(x+y)=expb(x)*expb(y)
    expb(xy)= [expb(x)]y

  • Logarithms are the inverse functions of exponential functions.
    expb(logb(x))=x
    logb(expb(x))=x

    In formula representation, these are written as follows:
    blogb(x)=x
    logb(bx)=x
  • Whenever you see a formula with an exponential, say b3x-2, you should be able to think in both formula and function modes interchangeably.
    b3x-2 = b3x b-2 = (b3)x b-2
    b3x-2 = expb(3x-2)
    The first mode allows us to recognize that b3x-2t is actually of the form Aqx where A=b-2 and q=b3. The second mode is useful to remind us that we really have a composition when we need to compute a derivative or when dealing with inverse functions.
  • There is a special base e that is most important because the corresponding exponential function is its own derivative.
    exp'e(x)=expe(x)
    d/dx[ex]=ex

    The natural exponential is written without a base. The corresponding inverse function is called the natural logarithm, ln(x).
    exp(ln(x))=x
    ln(exp(x))=x

    In formula representation, these are written as follows:
    eln(x)=x
    ln(ex)=x
    The x in these formulas, as always, is a placeholder. So any number or formula could be used in place of x.
  • Any exponential can be written in terms of the natural exponential. The key is to use the properties of exponents.
    b = eln(b)
    bx = [eln(b)]x = eln(b) x
    Another way to think of this is using composition of inverse functions: exp(ln( ))
    bx = exp(ln(bx))
    ln(bx) = x ln(b)
    bx = exp(x ln(b)) = ex ln(b)

    That is, by replacing bx by eln(b) x, we can write any exponential A bx in the form A ekx where k is the number ln(b).
  • Derivatives of exponentials use the basic property that exp'(x) = exp(x), or d/dx[ex] = ex. Usually, this also requires the chain rule: d/dx[eu] = eu u'.
    d/dx[e2x] = e2x (2) = 2e2x (u = 2x)
    d/dx[e-3x] = e-3x (-3) = -3e-3x (u = -3x)
    d/dt[e(t2-5t)] = e(t2-5t) (2t-5) = (2t-5)e(t2-5t) (u = t2-5t)

So this was a moderately long list. Perhaps just re-reading it helped you understand something better. Or perhaps you realize something is still confusing. Please post comments explaining why you are finding problems challenging or perhaps explaining what helped you suddenly understand what you had missed earlier.

Tuesday, November 17, 2009

Concerns about Derivatives

Okay, I have just finished grading HW #4. This assignment had some differentiation rules. But many of you are not comprehending the purpose of a derivative rule.

For example, we had the special rule for squares of functions:

d/dx[f2(x)] = 2 f(x) f '(x)

This means that any time there is a formula squared and you need to take its derivative, you can apply this rule.

(2x+3)2 corresponds to the function f(x)=2x+3 being squared. So since f '(x) = 2:

d/dx[(2x+3)2] = 2 (2x+3) (2) = 4(2x+3)

Similarly, (x2-4x+5)2 corresponds to the function f(x) = x2-4x+5 being squared, and f '(x)=2x-4

d/dx[(x2-4x+5)2] = 2 (x2-4x+5) (2x-4)

You need to be an expert at identifying the form of an expression in order to apply appropriate rules of differentiation, and next semester, rules of integration (anti-differentiation).