<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
	<channel>
		<title>Data Analytics Tools on Dataprd.Com</title>
		<link>https://dataprd.com/categories/data-analytics-tools/</link>
		<description>Recent content in Data Analytics Tools on Dataprd.Com</description>
		<generator>Hugo</generator>
		<language>en-us</language>
		
		
		
		
			<lastBuildDate>Sat, 15 Aug 2026 09:00:00 +0000</lastBuildDate>
		
			<atom:link href="https://dataprd.com/categories/data-analytics-tools/index.xml" rel="self" type="application/rss+xml" />
			<item>
				<title>Turning Natural Language Pricing into Structured Data for Automated Pricing</title>
				<link>https://dataprd.com/posts/llm-price-list-extraction/</link>
				<pubDate>Sat, 15 Aug 2026 09:00:00 +0000</pubDate>
				<guid>https://dataprd.com/posts/llm-price-list-extraction/</guid>
				<description>&lt;p&gt;Every accommodation booking system eventually runs into the same wall. Owners describe their prices in prose, and software needs rows in a table.&lt;/p&gt;&#xA;&lt;p&gt;What arrives looks like this, give or take a language. The colours show which sentence becomes which kind of row:&lt;/p&gt;&#xA;&lt;div class=&#34;owner-text&#34;&gt;&#xA;&lt;div class=&#34;owner-text-label&#34;&gt;Owner&#39;s text&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-season&#34;&gt;High season from 15 June to 22 August, 35,000 per night for up to 6 guests, extra person 5,000 per night.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-season&#34;&gt;Cleaning 15,000, free from 7 nights.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-holiday&#34;&gt;Christmas package 24 to 28 December, 4 nights 180,000.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-pet&#34;&gt;Dogs 2,000 per night, maximum two, up to 15 kg.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-discount&#34;&gt;Booking of 7 nights gets 1 night free.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;div class=&#34;owner-line&#34;&gt;&lt;span class=&#34;hl hl-rule&#34;&gt;Saturday to Saturday in high season.&lt;/span&gt;&lt;/div&gt;&#xA;&lt;/div&gt;&#xA;&lt;table&gt;&#xA;&#x9;&lt;thead&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Row&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;th&gt;Captures&lt;/th&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/thead&gt;&#xA;&#x9;&lt;tbody&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;span class=&#34;tag tag-season&#34;&gt;season&lt;/span&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Date range, nightly rate, guests included, extra guest fee, cleaning fee and when it is waived&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;span class=&#34;tag tag-holiday&#34;&gt;holiday&lt;/span&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Fixed night count and a package total&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;span class=&#34;tag tag-pet&#34;&gt;pet fee&lt;/span&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Per-animal rate, maximum animals, maximum weight&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;span class=&#34;tag tag-discount&#34;&gt;discount&lt;/span&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Type &lt;code&gt;night_bonus&lt;/code&gt;: 7 nights required, 1 night free&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&#x9;&#x9;&lt;tr&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;&lt;span class=&#34;tag tag-rule&#34;&gt;rule&lt;/span&gt;&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&#x9;&#x9;&lt;td&gt;Arrival and departure day, minimum stay&lt;/td&gt;&#xA;&#x9;&#x9;&#x9;&lt;/tr&gt;&#xA;&#x9;&lt;/tbody&gt;&#xA;&lt;/table&gt;&#xA;&lt;p&gt;Six sentences become five rows, with every ambiguity resolved on the way. The model returns this:&lt;/p&gt;</description>
			</item>
			<item>
				<title>Configure Apache Kylin with ODBC to work with MS PowerBI</title>
				<link>https://dataprd.com/posts/configure-apache-kylin-with-odbc-to-wowk-with-ms-powerbi/</link>
				<pubDate>Mon, 09 Jan 2017 21:13:16 +0000</pubDate>
				<guid>https://dataprd.com/posts/configure-apache-kylin-with-odbc-to-wowk-with-ms-powerbi/</guid>
				<description>&lt;h1 id=&#34;powerbi-and-kylin---reporting-from-hadoop-via-odbc&#34;&gt;PowerBI and Kylin - reporting from Hadoop via ODBC&lt;/h1&gt;&#xA;&lt;p&gt;This article discusses how to set up an ODBC interface for Kylin to work with Microsoft PowerBI. See previous article on the &lt;a href=&#34;https://dataprd.com/posts/evaluation-of-apache-kylin-1-5-4-1-with-hdp-2-5-performance-comparison-w-hive/&#34;&gt;detailed dataset, environment setup, on what Kylin is and how to create a cube in Kylin&lt;/a&gt;. For the tutorial&amp;rsquo;s purposes, we will analyze the previously loaded flight delay data with Hadoop, Hive, HBase, Kylin, Kylin ODBC connector and MS PowerBI as an interface. With &lt;a href=&#34;https://powerbi.microsoft.com/en-us/&#34;&gt;PowerBI&lt;/a&gt;, Microsoft provides a capable and simple BI tool for free (desktop version) - as competition heats up, this is a great strategy to gain some market share.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Evaluation of Apache Kylin 1.5.4.1 with HDP 2.5, performance comparison w Hive</title>
				<link>https://dataprd.com/posts/evaluation-of-apache-kylin-1-5-4-1-with-hdp-2-5-performance-comparison-w-hive/</link>
				<pubDate>Mon, 09 Jan 2017 20:24:45 +0000</pubDate>
				<guid>https://dataprd.com/posts/evaluation-of-apache-kylin-1-5-4-1-with-hdp-2-5-performance-comparison-w-hive/</guid>
				<description>&lt;p&gt;&lt;strong&gt;Apache Kylin&lt;/strong&gt; is a data cube solution on top of Hadoop providing an ODBC interface for BI tools. OLAP cubes boost performance for analytics via using a subset of data, enriched with pre-calculations on specific dimensions of interest. It enables loading dimensions from a Hive data source, therefore accelerating BI tool access via pre-calculating data and adding it to HBase. In our example we have a large dataset of flight information:&lt;/p&gt;</description>
			</item>
			<item>
				<title>Create a Hadoop Cluster easily by using PXE boot, Kickstart, Puppet and Ambari to auto-deploy nodes</title>
				<link>https://dataprd.com/posts/create-hadoop-cluster-easily-using-pxe-boot-kickstart-puppet-ambari-auto-deploy-nodes/</link>
				<pubDate>Fri, 08 Aug 2014 17:12:52 +0000</pubDate>
				<guid>https://dataprd.com/posts/create-hadoop-cluster-easily-using-pxe-boot-kickstart-puppet-ambari-auto-deploy-nodes/</guid>
				<description>&lt;p&gt;This tutorial is to showcase unattended and automatic install of multiple &lt;strong&gt;CentOS 6.5 x86_64 Hadoop nodes&lt;/strong&gt; pre-configured with &lt;strong&gt;Ambari-agents&lt;/strong&gt; and an &lt;strong&gt;Ambari-server&lt;/strong&gt; host. After configuring automatic install of bare metal (No OS pre-installed) nodes, deploying a Hadoop cluster will be a matter of clicks. The setup uses:&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;strong&gt;PXE boot&lt;/strong&gt; (for automatic OS install)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;TFTP server&lt;/strong&gt; (for PXE network install image)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Apache server&lt;/strong&gt; (to serve the kickstart file for unattended install)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;DHCP server&lt;/strong&gt; (for assigning IP addresses for the nodes)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;DNS server&lt;/strong&gt; (for internal domain name resolution)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Puppet-master&lt;/strong&gt; (for automatic configuration management of all hosts in the network, Ambari install included in Puppet manifests)&lt;/li&gt;&#xA;&lt;li&gt;&lt;strong&gt;Ambari-master and agents&lt;/strong&gt; (for managing Hadoop ecosystem deployment)&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;p&gt;The setup assumes that the nodes are on the &lt;em&gt;192.168.0.0/255.255.255.0&lt;/em&gt; network, the master is on &lt;em&gt;192.168.0.1&lt;/em&gt; and its hostname is &lt;em&gt;bigdata1.hdp&lt;/em&gt; The domain for the network server by the configuration is &lt;em&gt;hdp&lt;/em&gt; and the clients are named as &lt;em&gt;bigdata[1-254].hdp&lt;/em&gt; &lt;a href=&#34;https://dataprd.com/files/2014/08/provision.zip&#34; title=&#34;Hadoop Kickstart install&#34;&gt; Download all files here (configuration files, PXEBoot Linux image, Kickstart file and custom script for adding a node on the master).&lt;/a&gt;&lt;/p&gt;</description>
			</item>
			<item>
				<title>Analysis tutorial with Tableau Desktop</title>
				<link>https://dataprd.com/posts/analysis-tutorial-tableau-desktop/</link>
				<pubDate>Sun, 11 May 2014 17:09:31 +0000</pubDate>
				<guid>https://dataprd.com/posts/analysis-tutorial-tableau-desktop/</guid>
				<description>&lt;p&gt;Tableau Desktop supports visual analysis and data discovery, converts the raw information to easy to understand graphical format with interactive charts. No coding is required to create rich visualization. Tableau Business Intelligence toolset has a Desktop, Server and Cloud version (none open-source products but as good as worth a post on the open-bigdata blog). In this post I check its&lt;a href=&#34;http://www.tableausoftware.com/products/trial&#34;&gt; Desktop evaluation version&lt;/a&gt; that lets us connect to many data sources (including Hadoop, MySQL, Excel, Text, &amp;hellip;). I will use &lt;a href=&#34;https://dataprd.com/media/Weather.zip&#34;&gt;the same Weather.Csv&lt;/a&gt; as in the &lt;a href=&#34;https://dataprd.com/posts/analysis-fundamentals-tutorial/&#34;&gt;Hadoop analysis tutorial&lt;/a&gt;.&lt;/p&gt;</description>
			</item>
			<item>
				<title>Environment setup for big data analytics</title>
				<link>https://dataprd.com/posts/environment-setup-big-data-analytics/</link>
				<pubDate>Tue, 04 Mar 2014 15:18:50 +0000</pubDate>
				<guid>https://dataprd.com/posts/environment-setup-big-data-analytics/</guid>
				<description>&lt;p&gt;This article covers basic tools and technologies to use when conducting the first steps on big data analysis.&lt;/p&gt;&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Linux as the base OS&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://www.debian.org/&#34; title=&#34;Debian Linux&#34;&gt;Debian&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.ubuntu.com/&#34; title=&#34;Ubuntu Linux&#34;&gt;Ubuntu&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.redhat.com/&#34; title=&#34;Redhat Linux&#34;&gt;RedHat&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.centos.org/&#34; title=&#34;CentOS Linux&#34;&gt;CentOS&lt;/a&gt;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;li&gt;For basic data processing:&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Bash shell: environment for running multiple command-line Linux tools for data manipulation&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Comes pre-installed with Linux; check &lt;a href=&#34;http://www.gnu.org/software/bash/manual/bashref.html&#34; title=&#34;Bash reference manual&#34;&gt;Reference manual&lt;/a&gt; for usage&lt;/li&gt;&#xA;&lt;li&gt;Bash might be not the default Linux shell, &lt;a href=&#34;http://unix.stackexchange.com/questions/1373/how-do-i-switch-from-an-unknown-shell-to-bash&#34; title=&#34;Switch to Bash shell&#34;&gt;see how to switch to it&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;Learn Bash by &lt;a href=&#34;http://linuxconfig.org/bash-scripting-tutorial&#34; title=&#34;Bash examples&#34;&gt;examples&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;Most important Linux Commands and phenomena to master for data manipulation:&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.unix.com/man-page/linux/0/cat/&#34; title=&#34;Cat Reference Card&#34;&gt;cat&lt;/a&gt; - output/concatenate a stream / file&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.unix.com/man-page/POSIX/1/head/&#34; title=&#34;Head Reference Card&#34;&gt;head&lt;/a&gt;, &lt;a href=&#34;http://www.unix.com/man-page/freebsd/1/tail/&#34; title=&#34;Tail Reference Card&#34;&gt;tail&lt;/a&gt; - show the head or the tail of a stream / file&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.unix.com/man-page/OpenSolaris/1/grep&#34; title=&#34;Grep Reference Card&#34;&gt;grep&lt;/a&gt; - to filter streams / lines of a file&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.unix.com/man-page/freebsd/1/sed/&#34; title=&#34;Sed Reference Card&#34;&gt;sed&lt;/a&gt; - to manipulate streams / lines of a file&lt;/li&gt;&#xA;&lt;li&gt;Regular expressions - used in a wide set of Linux tools&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.grymoire.com/unix/Regular.html&#34; title=&#34;Explanation and tutorials&#34;&gt;Explanation and tutorials&lt;/a&gt;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;li&gt;AWK - simple data reformatter with compact coding features&#xA;&lt;ul&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.gnu.org/software/gawk/manual/gawk.html&#34; title=&#34;The GNU Awk User&#39;s Guide&#34;&gt;The GNU Awk User&amp;rsquo;s Guide&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.grymoire.com/Unix/Awk.html&#34; title=&#34;AWK Tutorials&#34;&gt;AWK Tutorials&lt;/a&gt;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;li&gt;Python - easy to learn, effective programming language with a huge amount of libraries available for various tasks. Great for data manipulation used from the command line.&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Python &lt;a href=&#34;http://www.python.org/&#34; title=&#34;Python Download&#34;&gt;Download&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.python.org/doc/&#34; title=&#34;Python Documentation&#34;&gt;Python Documentation&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;https://wiki.python.org/moin/BeginnersGuide&#34; title=&#34;Python Beginners Guide&#34;&gt;Python Beginners Guide&lt;/a&gt;&lt;/li&gt;&#xA;&lt;li&gt;&lt;a href=&#34;http://www.learnpython.org/&#34; title=&#34;Python Tutorials&#34;&gt;Python Tutorials&lt;/a&gt;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;li&gt;And the big data analysis framework chosen based on the type of data analyzed. For the first step tutorials our suggestion would be:&#xA;&lt;ul&gt;&#xA;&lt;li&gt;Hadoop, single cluster setup (can be downloaded pre-installed to a virtual appliance)&lt;/li&gt;&#xA;&lt;li&gt;Java based MapReduce programs&lt;/li&gt;&#xA;&lt;li&gt;Pig MapReduce query language&lt;/li&gt;&#xA;&lt;/ul&gt;&#xA;&lt;/li&gt;&#xA;&lt;/ul&gt;</description>
			</item>
	</channel>
</rss>
