
The EM Algorithm 193
and
log f
X
(x|θ) = log f
X,Y
(x, y|θ) − log f
Y |X
(y|x),
we find that maximizing E(log f
X
(x|θ)|y, θ
k
) is equivalent to minimizing
H(θ
k
,θ)=
f
X,Y
(x, y|θ
k
)logf
X,Y
(x, y|θ)dx. (13.31)
With f (θ)=f
Y
(y|θ), and b(θ)=f
X,Y
(x, y|θ), this problem fits the frame-
work of the non-stochastic EM algorithm and is equivalent to minimizing
G(θ
k
,θ)=KL(b(θ
k
),b(θ)) −f (θ).
Once again, we may conclude that the likelihood function is non-decreasing
and that the sequence {KL(b(θ
k
),b(θ
k+1
))} converges to zero.
In the discrete case in which Y = h(X) the conditional probability
f
Y |X
(y|x, θ)isδ(y −h(x)), as a function of y, for given x, and is the char-
acteristic