Note that the dual variables associated with LP constraints 1, 2, and 3, respectively, are
identical to the expected total discounted rewards,
5,563.74, 5, 704.86, 6,020.57vv v
, respectively, obtained by policy iteration in
5.5g) Begin policy iteration with
0.9
D
by choosing an initial policy given by the
vector
[1 1 1] [RP RP PM]
TT
d
. The initial decision vector
d
, along with the
associated transition probability matrix
P
and the cost vector
q
, are shown below.
1 0.1 0.4 0.5 300
ªº ª º ª º
¬¼ ¬ ¼ ¬ ¼
The VDEs are
1123
300 0.9(0.1 0.4 0.5 )
vvvv
The solution of the VDEs is
12 3
5,563.74, 5, 704.86, 6,020.57vv v
First policy improvement for machine maintenance model
State Decision Test Quantity
11 2 2 3 3
Test Quantity
()
kkkk
iiii
qpvpvpv
D