Tuesday, June 21, 2016

How to list word occurences using CountVectorizer from Scikit Learn

A simple and efficient way to get document frequency counts of words from a corpus is to use CountVectorizer from Scikit Learn

Getting back to the word from the index is not immediately obvious, here's how to do it:

import numpy as np
from sklearn.feature_extraction.text import CountVectorizer

docs = <load your docs as an iterable>

count_vect = CountVectorizer()
doc_counts = count_vect.fit_transform(docs)  # this is of type scipy.sparse.csr.csr_matrix which is why we need to use

.ravel() below.

word_counts = zip(count_vect.get_feature_names(), np.asarray(doc_counts.sum(axis=0)).ravel())
word_counts = sorted(word_counts, key=lambda idx: -1 * idx[1] )
 

# Display top 100 words by frequency
word_counts[:100]

Thursday, June 9, 2016

How to 7zip each file seperately in a directory (Windows)

Let's zip up all those log files sitting in that directory into seperate 7z files....

FOR %i IN (*.*) DO "C:\Program Files\7-Zip\7z.exe" a -mx=9 "%i.7z" "%i"

Refs:
http://superuser.com/questions/312652/how-do-i-create-seperate-7z-files-from-each-selected-directory-with-7zip-command
http://askubuntu.com/questions/491223/7z-ultra-settings-for-zip-format

Monday, April 18, 2016

Managing Java versions and Eclipse

Some little bits on managing Java versions and Eclipse:

Managing JRE installations in Eclipse
http://www.codejava.net/ides/eclipse/managing-jre-installations-in-eclipse

How to switch JDK version on Mac OS X
https://www.jayway.com/2014/01/15/how-to-switch-jdk-version-on-mac-os-x-maverick/
Edit your ~/.bash_profile and add the following:
function setjdk() {
  if [ $# -ne 0 ]; then
   removeFromPath '/System/Library/Frameworks/JavaVM.framework/Home/bin'
   if [ -n "${JAVA_HOME+x}" ]; then
    removeFromPath $JAVA_HOME
   fi
   export JAVA_HOME=`/usr/libexec/java_home -v $@`
   export PATH=$JAVA_HOME/bin:$PATH
  fi
 }
 function removeFromPath() {
  export PATH=$(echo $PATH | sed -E -e "s;:$1;;" -e "s;$1:?;;")
 }
setjdk 1.7


Monday, March 21, 2016

Installing Java 8 on Ubuntu

Quick, easy point of reference for installing Java 8 on Ubuntu:

http://stackoverflow.com/questions/25549492/install-jdk8-in-ubuntu-14-04

sudo add-apt-repository ppa:webupd8team/java
sudo apt-get update
sudo apt-get install oracle-java8-installer


http://askubuntu.com/questions/315646/update-java-alternatives-vs-update-alternatives-config-java
sudo update-alternatives --config java


Tuesday, February 23, 2016

Debugging spark job incorrect JVM jar file loaded

I recently had an issue with incorrect version of a JVM jar file being loaded in a Spark job on Hadoop.

This Scala code snippet helped debug the issue and determine the loaded jar file path:

val jarPath = classOf[MyObject].getProtectionDomain().getCodeSource().getLocation().getPath()

In my case this pointed to an old version of guava:
/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-11.0.jar

Then by setting spark.driver.extraClassPath and spark.executor.extraClassPath arguments for spark-submit, the correct version of the jar file was loaded successfully:

spark-submit --class com.MyClass <other_spark_args> --conf "spark.driver.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" --conf "spark.executor.extraClassPath=/opt/cloudera/parcels/CDH-5.5.1-1.cdh5.5.1.p0.11/jars/guava-15.0.jar" /home/mypath/myjarfile.jar <my_job_params>

For more info with extraClassPath, see: http://spark.apache.org/docs/latest/configuration.html

Sunday, November 1, 2015

Convert binary word2vec model to text vectors

If you have a binary model generated from google's awesome and super fast word2vec word embeddings tool, you can easily use python with gensim to convert this to a text representation of the word vectors.

Input: binary word embedding model from google's word2vec tool

Output: text vectors for word embeddings

Python conversion code:
from gensim.models import word2vec
model = word2vec.Word2Vec.load_word2vec_format('path/to/mymodel.bin', binary=True)
model.save_word2vec_format('path/to/
mymodel.txt', binary=False)

I recommend using Anaconda from Continuum Analytics for a bundled python distribution.  To install gensim in Anaconda just type: conda install gensim :)


Original ref: https://www.kaggle.com/c/word2vec-nlp-tutorial/forums/t/13828/how-to-convert-bin-file-of-word2vec-model-into-txt-r/91564

Thursday, October 22, 2015

Analyse consecutive timeseries pairs in Spark

With time series data it is useful to construct pairs of elements for analysis.

Here is a way to construct a Spark RDD from a time series that has the pairs together in the final RDD:


val arr = Array((1, "A"), (8, "D"), (7, "C"), (3, "B"), (9, "E"))
val rdd = sc.parallelize(arr)
val sorted = rdd.sortByKey(true)
val zipped = sorted.zipWithIndex.map(x => (x._2, x._1))
val pairs = zipped.join(zipped.map(x => (x._1 - 1, x._2))).sortBy(_._1)


Which produces the consecutive elements as pairs in the RDD for further processing:
(0,((1,A),(3,B)))
(1,((3,B),(7,C)))
(2,((7,C),(8,D)))
(3,((8,D),(9,E)))


Ref: https://www.mail-archive.com/user@spark.apache.org/msg39353.html