HIVE精炼笔记总结—

1.1. 数据类型

1.1.1. 数字类型

TINYINT (1-bytesigned integer, from -128 to 127)

SMALLINT (2-bytesigned integer, from -32,768 to 32,767)

INT/INTEGER(4-byte signed integer, from -2,147,483,648 to 2,147,483,647)

BIGINT (8-bytesigned integer,from -9,223,372,036,854,775,808 to 9,223,372,036,854,775,807)

FLOAT (4-byte single precisionfloating point number)

DOUBLE (8-byte double precisionfloating point number)

示例：

create table t_test(a string,b int,c bigint,d float,e double,f tinyint,g smallint)

1.1.2. 日期时间类型

TIMESTAMP (Note:Only available starting with Hive 0.8.0)

DATE (Note:Only available starting with Hive 0.12.0)

示例，假如有以下数据文件：

1,zhangsan,1985-06-30

2,lisi,1986-07-10

3,wangwu,1985-08-09

那么，就可以建一个表来对数据进行映射

create table t_customer(id int,namestring,birthday date)

row format delimited fields terminated by',';

然后导入数据

load data local inpath '/root/customer.dat'into table t_customer;

然后，就可以正确查询

1.1.3. 字符串类型

STRING

VARCHAR (Note:Only available starting with Hive 0.12.0)

CHAR (Note:Only available starting with Hive 0.13.0)

1.1.4. 混杂类型

BOOLEAN

BINARY (Note: Only available startingwith Hive 0.8.0)

1.1.5. 复合类型

1.1.5.1. array数组类型

arrays: ARRAY<data_type> (Note:negative values and non-constant expressions are allowed as of Hive 0.14.)

示例：array类型的应用

假如有如下数据需要用hive的表去映射：

钢铁侠3,唐尼:小辣椒:哈皮,2013-05-04

蜘蛛侠英雄归来,荷兰弟:唐尼,2017-07-07

设想：如果主演信息用一个数组来映射比较方便

建表：

create tablet_movie(moive_name string,actors array<string>,first_show date)

row formatdelimited fields terminated by ','

collection itemsterminated by ':';

导入数据：

load data local inpath '/root/movie.dat'into table t_movie;

查询：

select * from t_movie;

select moive_name,actors[0] from t_movie;

select moive_name,actors from t_movie where array_contains(actors,'唐尼');

select moive_name,size(actors) fromt_movie;

1.1.5.2. map类型

maps: MAP<primitive_type,data_type> (Note: negative values and non-constant expressions areallowed as of Hive 0.14.)

1) 假如有以下数据：

1,zhangsan,father:xiaoming#mother:xiaohuang#brother:xiaoxu,28

2,lisi,father:mayun#mother:huangyi#brother:guanyu,22

3,wangwu,father:wangjianlin#mother:ruhua#sister:jingtian,29

4,mayun,father:mayongzhen#mother:angelababy,26

可以用一个map类型来对上述数据中的家庭成员进行描述

2) 建表语句：

create table t_person(id int,name string,family_members map<string,string>,age int)

row format delimited fields terminated by','

collection items terminated by '#'

map keys terminated by':';

3) 查询

select * from t_person;

## 取map字段的指定key的值

select id,name,family_members['father'] asfather from t_person;

## 取map字段的所有key

select id,name,map_keys(family_members) asrelation from t_person;

## 取map字段的所有value

select id,name,map_values(family_members)from t_person;

select id,name,map_values(family_members)[0]from t_person;

## 综合：查询有brother的用户信息

select id,name,father

from

(select id,name,family_members['brother'] as father from t_person) tmp

where father is not null;

1.1.5.3. struct类型

structs: STRUCT<col_name :data_type, ...>

1) 假如有如下数据：

1,zhangsan,18:male:beijing

2,lisi,28:female:shanghai

其中的用户信息包含：年龄：整数，性别：字符串，地址：字符串

设想用一个字段来描述整个用户信息，可以采用struct

2) 建表：

create table t_person_struct(id int,namestring,info struct<age:int,sex:string,addr:string>)

row format delimited fields terminatedby ','

collection items terminated by ':';

3) 查询

select * from t_person_struct;

select id,name,info.agefrom t_person_struct;

1.2. 修改表定义

仅修改Hive元数据，不会触动表中的数据，用户需要确定实际的数据布局符合元数据的定义。

修改表名：

ALTER TABLE table_name RENAME TOnew_table_name

示例：alter table t_1 rename to t_x;

修改分区名：

alter table t_partitionpartition(department='xiangsheng',sex='male',howold=20) rename topartition(department='1',sex='1',howold=20);

添加分区：

alter table t_partition addpartition (department='2',sex='0',howold=40);

删除分区：

alter table t_partition droppartition (department='2',sex='2',howold=24);

修改表的文件格式定义：

ALTER TABLE table_name [PARTITIONpartitionSpec] SET FILEFORMAT file_format

alter table t_partitionpartition(department='2',sex='0',howold=40 ) set fileformat sequencefile;

修改列名定义：

ALTER TABLE table_name CHANGE[COLUMN] col_old_name col_new_name column_type [COMMENTcol_comment][FIRST|(AFTER column_name)]

alter table t_user change price jiage floatfirst;

增加/替换列：

ALTER TABLE table_nameADD|REPLACE COLUMNS (col_name data_type[COMMENT col_comment], ...)

alter table t_user add columns (sexstring,addr string);

alter table t_user replace columns (idstring,age int,price float);