Unlock access to all the studying documents.
View Full Document
CHAPTER 4 A MARKOV CHAIN WITH REWARDS (MCR)
4.1
4.1 a)
(3) 255.15pq
4.1 b)
S
[0.209651 0.480865 0.134775 0.174709]
4.1 c)
1.996806
5.58754
1.011652
3
1.905627
4.41916
2.335321
1.33989
4
1.750339
3.99176
1.464695
2.79321
4.2
4.2 a)
State 5 4 1 2 3
5, Sell to son 10000
4, Sell to daughter 0 1 0 0 0 0
1, Weekly income $10,000 0.05 0.10 0.40 0.25 0.20
I
PDQ
ªº
«
¬¼
»
º
ª
R
q
000,20$3
$15,000 2
daughter toSell 4
son toSell 5
incomeWeekly State
»
º
«
ª
»
º
«
ª
000,10$
1
1
1
q
4.2 b)
4.2 c)
4.2 d) Given that the firm’s income in the first week is $15,000, the expected number of
absorption in state 4. This quantity is the sum of the entries in row 2 of the fundamental
1
ij
iC
410 0 0
0.1(1) 0.4(0.569) 0.25(0.561) 0.2(0.506)
10.569 0.569 0.569 0.569
3
ªº
«»
«»
41 2 3
4123
41 0 0 0
0
1 0.176 0.4 0.246 0.178
I
ªº
«»
ªº
¬¼
1
()UIQ
4.2 e)
2
T
Suppose the following additional question is asked.
income that will be earned before the firm is sold to the daughter?
T
Uq
Given that the firm’s income in the first week is $15,000, the expected total income that
will be earned before the firm is sold to the daughter =
2
()
T
Uq
50341.8.
4. 3
4.3a)
The components
i
q
of the cost vector are
60$57$)10/3(10$
0
q
2
20$10$)1(10$
3
q
3210
4.3b)
State 0 1 2 3 State
230
21
200
33
320
31000
4.3c) In matrix form the VDEs are
0 0
3 3
60 0.30 0.70 0 0
20 1 0 0 0
After setting 0,the solution for the average cost per year and
vv
g
vv
g
v
ªº ªº
ªº ª º ª º
«» «»
«» « » « »
¬¼ ¬ ¼ ¬ ¼
¬¼ ¬¼
the relative costs is
Thus, if this Markov chain with costs operates over an infinite planning horizon, the
4.3e) Make the target recurrent state 0 an absorbing state.
º
ª
»
»
º
«
«
ª
01
04286.005714.0
0001
1
0
º
ª
A
q
State
600
4.3f) With discounting, the model has the following transition probability matrix, P,
discount factor
9.0
D
, discounted matrix,
P
D
, and cost vector, q.
20
30
40
60
0009.0
3.0006.0
03857.005143.0
0063.027.0
3
2
1
0
0001
3333.0.006667.0
04286.005714.0
0070.030.0
3
2
1
0
»
»
»
»
¼
º
«
«
«
«
¬
ª
»
»
»
»
¼
º
«
«
«
«
¬
ª
»
»
»
»
¼
º
«
«
«
«
¬
ª
In matrix form the VDEs are
Thus, if this Markov chain with discounted costs operates over an infinite planning horizon,
4.3g)
>@
0010
)0(
p
)0()1(
At the end of year 1, the expected cost of operation is
)1(
20
30
40
60
»
»
»
»
¼
º
«
«
«
«
¬
ª
)1()2(
At the end of year 2, the expected cost of operation is
)2(
40
60
»
»
º
«
«
ª
At the end of year 3, the expected cost of operation is
)3(
20
30
40
60
»
»
»
»
¼
º
«
«
«
«
¬
ª
The expected total cost of operation after three years is
4.3h) Since value iteration assumes zero terminal costs at the end of a planning horizon, it
will be executed over four years to calculate the expected total cost after three years.
v1(0) = 181.89
4.4)
4.4a)
State $0 $5,000 $1,000 $2,000 $3,000 $4,000
$0100000
$1,000 0.6 0 0 0.4 0 0
$4,000 0 0.4 0 0 0.6 0
P
,
0
5000
3000
4000
ªº
«»
«»
«»
¬¼
>@
001000
)0(
p
)0()1(
State $0 $5,000 $1,000 $2,000 $3,000 $4,000
$0100000
$2,000 0 0 0.6 0 0.4 0
>@
PPpp 04.006.000
)1()2(
State $0 $5,000 $1,000 $2,000 $3,000 $4,000
$0100000
$3,000 0 0 0 0.6 0 0.4
4000
3000
2000
»
»
»
»
¼
«
«
«
«
¬
4.4b) Value iteration is modified by setting v(2) = q, and making the reward vector a null vector.
When an absorbing state is entered, the expected total reward does not change because the
game has ended.
-1
4.5
State 5 4 1 2 3
5, Scrapped 10000
3, Stage 3 0.04 0.90 0 0 0.06
State Cost (in dollars)
560
3500
4.5b) The expected cost per item started is
1312111415 500$400$300$20$60$100$uuuff
The expected total cost of an item sold is equal to
141312111415
/)500$400$300$20$60$100($ fuuuff
The expected total cost of an item sold is equal to
(Uq)1 = 1192.185