计算TensorFlow中两个输入集合的每对之间的成对距离

Question

I have two collections. 我有两个收藏。 One consists of m ₁ points in k dimensions and another one of m ₂ points in k dimensions. 一个由米在K个维度₁分和K个维度的米₂分另一个。 I need to calculate pairwise distance between each pair of the two collections. 我需要计算两个集合的每对之间的成对距离。

Basically having two matrices A _{m ₁ , k} and B _{m ₂ , k} I need to get a matrix C _{m ₁ , m ₂} . 基本上有两个矩阵A _{m ₁ ，k}和B _{m ₂ ，k} I需要得到一个矩阵C _{m ₁ ，m ₂} 。

I can easily do this in scipy by using distance.sdist and select one of many distance metrics, and I also can do this in TF in a loop, but I can't figure out how to do this with matrix manipulations even for Eucledian distance. 我可以通过使用distance.sdist轻松地在scipy中执行此操作，并选择许多距离度量之一，而且我也可以在TF中循环执行此操作，但是即使对于Eucledian距离，我也无法弄清楚如何使用矩阵操作来执行此操作。

Answer 1

After a few hours I finally found how to do this in Tensorflow. 几个小时后，我终于在Tensorflow中找到了如何进行此操作。 My solution works only for Eucledian distance and is pretty verbose. 我的解决方案仅适用于Eucledian距离，并且非常冗长。 I also do not have a mathematical proof (just a lot of handwaving, which I hope to make more rigorous): 我也没有数学上的证明（只是做了大量的手工操作，我希望使其更加严格）：

import tensorflow as tf
import numpy as np
from scipy.spatial.distance import cdist

M1, M2, K = 3, 4, 2

# Scipy calculation
a = np.random.rand(M1, K).astype(np.float32)
b = np.random.rand(M2, K).astype(np.float32)
print cdist(a, b, 'euclidean'), '\n'

# TF calculation
A = tf.Variable(a)
B = tf.Variable(b)

p1 = tf.matmul(
    tf.expand_dims(tf.reduce_sum(tf.square(A), 1), 1),
    tf.ones(shape=(1, M2))
)
p2 = tf.transpose(tf.matmul(
    tf.reshape(tf.reduce_sum(tf.square(B), 1), shape=[-1, 1]),
    tf.ones(shape=(M1, 1)),
    transpose_b=True
))

res = tf.sqrt(tf.add(p1, p2) - 2 * tf.matmul(A, B, transpose_b=True))

with tf.Session() as sess:
    sess.run(tf.global_variables_initializer())
    print sess.run(res)

Answer 2

This will do it for tensors of arbitrary dimensionality (ie containing (..., N, d) vectors). 这将对任意维度的张量（即包含（...，N，d）向量）进行处理。 Note that it isn't between collections (ie not like scipy.spatial.distance.cdist ) it's instead within a single batch of vectors (ie like scipy.spatial.distance.pdist ) 请注意，它不在集合之间（即不像scipy.spatial.distance.cdist ），而是在单批向量内（例如scipy.spatial.distance.pdist ）

import tensorflow as tf
import string

def pdist(arr):
    """Pairwise Euclidean distances between vectors contained at the back of tensors.

    Uses expansion: (x - y)^T (x - y) = x^Tx - 2x^Ty + y^Ty 

    :param arr: (..., N, d) tensor
    :returns: (..., N, N) tensor of pairwise distances between vectors in the second-to-last dim.
    :rtype: tf.Tensor

    """
    shape = tuple(arr.get_shape().as_list())
    rank_ = len(shape)
    N, d = shape[-2:]

    # Build a prefix from the array without the indices we'll use later.
    pref = string.ascii_lowercase[:rank_ - 2]

    # Outer product of points (..., N, N)
    xxT = tf.einsum('{0}ni,{0}mi->{0}nm'.format(pref), arr, arr)

    # Inner product of points. (..., N)
    xTx = tf.einsum('{0}ni,{0}ni->{0}n'.format(pref), arr, arr)

    # (..., N, N) inner products tiled.
    xTx_tile = tf.tile(xTx[..., None], (1,) * (rank_ - 1) + (N,))

    # Build the permuter. (sigh, no tf.swapaxes yet)
    permute = list(range(rank_))
    permute[-2], permute[-1] = permute[-1], permute[-2]

    # dists = (x^Tx - 2x^Ty + y^Tx)^(1/2). Note the axis swapping is necessary to 'pair' x^Tx and y^Ty
    return tf.sqrt(xTx_tile - 2 * xxT + tf.transpose(xTx_tile, permute))

计算TensorFlow中两个输入集合的每对之间的成对距离

问题描述

2 个解决方案

解决方案1
3 2017-05-08 04:08:24

解决方案2
0 2018-08-26 04:48:59

计算TensorFlow中两个输入集合的每对之间的成对距离

问题描述

2 个解决方案

解决方案1 3 2017-05-08 04:08:24

解决方案2 0 2018-08-26 04:48:59

解决方案1
3 2017-05-08 04:08:24

解决方案2
0 2018-08-26 04:48:59