Hive表的基本操作

柯廣發表於2021-01-10

原文網址 : https://www.cnblogs.com/data-magnifier/p/14259325.html

1. 建立表

create table語句遵從sql語法習慣，只不過Hive的語法更靈活。例如，可以定義表的資料檔案儲存位置，使用的儲存格式等。

create table if not exists test.user1(
name string comment 'name',
salary float comment 'salary',
address struct<country:string, city:string> comment 'home address'
)
comment 'description of the table'
partitioned by (age int)
row format delimited fields terminated by '\t'
stored as orc;

沒有指定external關鍵字，則為管理表，跟mysql一樣，if not exists如果表存在則不做操作，否則則新建表。comment可以為其做註釋，分割槽為age年齡，列之間分隔符是\t，儲存格式為列式儲存orc，儲存位置為預設位置，即引數hive.metastore.warehouse.dir（預設：/user/hive/warehouse）指定的hdfs目錄。

2. 拷貝表

使用like可以拷貝一張跟原表結構一樣的空表，裡面是沒有資料的。

create table if not exists test.user2 like test.user1;

3. 檢視錶結構

通過desc [可選引數] tableName命令檢視錶結構，可以看出拷貝的表test.user1與原表test.user1的表結構是一樣的。

hive> desc test.user2;
OK
name                	string              	name                
salary              	float               	salary              
address             	struct<country:string,city:string>	home address        
age                 	int                 	                    
	 	 
# Partition Information	 	 
# col_name            	data_type           	comment             
	 	 
age                 	int

也可以加formatted，可以看到更加詳細和冗長的輸出資訊。

hive> desc formatted test.user2;
OK
# col_name            	data_type           	comment             
	 	 
name                	string              	name                
salary              	float               	salary              
address             	struct<country:string,city:string>	home address        
	 	 
# Partition Information	 	 
# col_name            	data_type           	comment             
	 	 
age                 	int                 	                    
	 	 
# Detailed Table Information	 	 
Database:           	test                	 
Owner:              	hdfs                	 
CreateTime:         	Mon Dec 21 16:37:57 CST 2020	 
LastAccessTime:     	UNKNOWN             	 
Retention:          	0                   	 
Location:           	hdfs://nameservice2/user/hive/warehouse/test.db/user2	 
Table Type:         	MANAGED_TABLE       	 
Table Parameters:	 	 
	COLUMN_STATS_ACCURATE	{\"BASIC_STATS\":\"true\"}
	numFiles            	0                   
	numPartitions       	0                   
	numRows             	0                   
	rawDataSize         	0                   
	totalSize           	0                   
	transient_lastDdlTime	1608539877          
	 	 
# Storage Information	 	 
SerDe Library:      	org.apache.hadoop.hive.ql.io.orc.OrcSerde	 
InputFormat:        	org.apache.hadoop.hive.ql.io.orc.OrcInputFormat	 
OutputFormat:       	org.apache.hadoop.hive.ql.io.orc.OrcOutputFormat	 
Compressed:         	No                  	 
Num Buckets:        	-1                  	 
Bucket Columns:     	[]                  	 
Sort Columns:       	[]                  	 
Storage Desc Params:	 	 
	field.delim         	\t                  
	serialization.format	\t

4. 刪除表

這跟sql中刪除命令drop table是一樣的：

drop table if exists table_name;

對於管理表（內部表），直接把表徹底刪除了；對於外部表，還需要刪除對應的hdfs檔案才會徹底將這張表刪除掉，為了安全，通常hadoop叢集是開啟回收站功能的，刪除外表表的資料就在回收站，後面如果想恢復也是可以恢復的，直接從回收站mv到hive對應目錄即可。

5. 修改表

大多數表屬性可以通過alter table來修改。

5.1 表重新命名

alter table test.user1 rename to test.user3;

5.2 增、修、刪分割槽

增加分割槽使用命令alter table table_name add partition(...) location hdfs_path

alter table test.user2 add if not exists
partition (age = 101) location '/user/hive/warehouse/test.db/user2/part-0000101'
partition (age = 102) location '/user/hive/warehouse/test.db/user2/part-0000102'

修改分割槽也是使用alter table ... set ...命令

alter table test.user2 partition (age = 101) set location '/user/hive/warehouse/test.db/user2/part-0000110'

刪除分割槽命令格式是alter table tableName drop if exists partition(...)

alter table test.user2 drop if exists partition(age = 101)

5.3 修改列資訊

可以對某個欄位進行重新命名，並修改位置、型別或者註釋：

修改前：

hive> desc user_log;
OK
userid              	string              	                    
time                	string              	                    
url                 	string

修改列名time為times，並且使用after把位置放到url之後，本來是在之前的。

alter table test.user_log
change column time times string
comment 'salaries'
after url;

再來看錶結構：

hive> desc user_log;
OK
userid              	string              	                    
url                 	string              	                    
times               	string              	salaries

time -> times，位置在url之後。

5.4 增加列

hive也是可以新增列的：

alter table test.user2 add columns (
birth date comment '生日',
hobby string comment '愛好'
);

5.5 刪除列

刪除列不是指定列刪除，需要把原有所有列寫一遍，要刪除的列排除掉即可：

hive> desc test.user3;
OK
name                	string              	name                
salary              	float               	salary              
address             	struct<country:string,city:string>	home address        
age                 	int                 	                    
	 	 
# Partition Information	 	 
# col_name            	data_type           	comment             
	 	 
age                 	int

如果要刪除列salary，只需要這樣寫：

alter table test.user3 replace columns(
name string,
address struct<country:string,city:string>
);

這裡會報錯：

FAILED: Execution Error, return code 1 from org.apache.hadoop.hive.ql.exec.DDLTask. Replacing columns cannot drop columns for table test.user3. SerDe may be incompatible

這張test.user3表是orc格式的，不支援刪除，如果是textfile格式，上面這種replace寫法是可以刪除列的。通常情況下不會輕易去刪除列的，增加列倒是常見。

5.6 修改表的屬性

可以增加附加的表屬性，或者修改屬性，但是無法刪除屬性：

alter table tableName set tblproperties(
    'key' = 'value'
);

舉例：這裡新建一張表：

create table t8(time string,country string,province string,city string)
row format delimited fields terminated by '#' 
lines terminated by '\n' 
stored as textfile;

這條語句將t8表中的欄位分隔符'#'修改成'\t';

alter table t8 set serdepropertyes('field.delim'='\t');

關注公眾號：Java大資料與資料倉儲，領取資料，學習大資料技術。

Hive的基本操作用法
2024-11-10
Hive
Hive學習之基本操作
2018-11-30
Hive
hadoop之旅8-centerOS7 : Hive的基本操作
2018-11-04
HadoopROSHive
Spark操作Hive分割槽表
2018-12-07
SparkHive
線性表的基本操作
2019-01-27
MySQL資料表的基本操作
2018-04-28
MySql
04 MySQL 表的基本操作-DDL
2024-04-03
MySql
openpyxl 操作 Excel表的格基本用法
2021-11-29
Excel
非常詳細地Hive的基本操作和一些注意事項
2020-11-25
Hive
HIVE基本語法以及HIVE分割槽
2018-09-20
Hive
Hive 常用操作
2018-08-20
Hive
SQL—對資料表內容的基本操作
2018-05-14
SQL
Hive高階操作-查詢操作
2024-06-28
Hive
hive建表
2024-10-05
Hive
Go 操作 Redis 的基本操作
2019-09-19
GoRedis
MySQL入門系列：資料庫和表的基本操作
2019-03-07
MySql資料庫
Hive的基本介紹以及常用函式
2020-06-04
Hive函式
hive04_DQL操作
2024-08-08
Hive
hive02_SQL操作
2024-07-26
HiveSQL
Docker的基本操作
2018-09-17
Docker
MySQL的基本操作
2018-06-04
MySql
git的基本操作
2018-03-07
Git
Oracle 操作表結構基本語法及示例
2018-07-12
Oracle
Go語言學習教程：xorm表基本操作及高階操作
2019-04-04
GoORM
資料庫表的基本操作和編碼格式設定
2020-10-03
資料庫
Hive 表的兩種分類
2020-11-30
Hive
[hive]hive資料模型中四種表
2018-08-14
Hive模型
Hive內部表和外部表的區別
2018-08-14
Hive
Hadoop實戰：Hive操作使用
2019-01-14
HadoopHive
hive03_高階操作
2024-07-26
Hive
JS — 物件的基本操作
2019-01-16
JS物件
Spring Boot的基本操作
2018-11-26
Spring Boot
Docker映象的基本操作
2018-05-15
Docker
Hbase shell的基本操作
2018-04-27
react的基本操作(1)
2018-07-30
React
Vim命令的基本操作
2020-10-10
Numpy的基本操作（五）
2020-10-04
陣列的基本操作
2024-08-25
陣列